--- license: agpl-3.0 pipeline_tag: image-segmentation library_name: ultralytics base_model: "Ultralytics/YOLO26" tags: - semantic-segmentation - aerial-imagery - drone - aeroscapes - yolo26 - ultralytics - pytorch - computer-vision datasets: - dronefreak/Aeroscapes metrics: - miou - pixel-accuracy --- # YOLO26s-sem Finetuned on AeroScapes ![License](https://img.shields.io/badge/License-AGPL--3.0-1f6feb?style=flat-square) ![Framework](https://img.shields.io/badge/Framework-Ultralytics-6a5acd?style=flat-square) ![Dataset](https://img.shields.io/badge/Dataset-AeroScapes-0aa1a7?style=flat-square) ![mIoU](https://img.shields.io/badge/mIoU-65.44%25-e8a33d?style=flat-square) ![Status](https://img.shields.io/badge/Status-Trained-2ea44f?style=flat-square) ![Maintained](https://img.shields.io/badge/Maintained-yes-17a2b8?style=flat-square) Fine-tuned YOLO26s semantic segmentation model for aerial drone imagery using the AeroScapes benchmark dataset. This model is part of the **AeroScapes Semantic Segmentation Model Zoo**, a collection of CABiNet and YOLO26 models trained and evaluated under a common pipeline for aerial semantic segmentation.

YOLO26s-sem on AeroScapes: Input / Ground Truth / Prediction

Qualitative results on AeroScapes val-split examples — single-scale (imgsz=896) inference, no TTA. --- ## Performance | Metric | Score | | ------------------- | --------------- | | mIoU | 65.44 | | Pixel Accuracy | 93.22 | | Parameters (M) | 6.50 | | FLOPs (GFLOPs @ 720px) | 21.9 | --- ## AeroScapes Model Zoo | Rank | Model | mIoU (%) | Pixel Acc (%) | Params (M) | FLOPs (GFLOPs) | | ---- | --------------------- | ------------- | ------------------ | ----------------- | ----------------- | | 1 | YOLO26x-sem | 68.36 | 93.91 | 40.15 | 213.0 | | 2 | YOLO26l-sem | 68.0 | 93.16 | 17.86 | 95.1 | | 3 | CABiNet (MobileNetV3-Large) | 67.83 | 94.32 | 9.18 | 27.4 | | 4 | YOLO26m-sem | 66.97 | 93.51 | 14.31 | 75.2 | | 5 | YOLO26s-sem | 65.44 | 93.22 | 6.50 | 21.9 | | 6 | YOLO26n-sem | 64.86 | 92.84 | 1.63 | 5.6 | --- ## Per-Class IoU (%) | Class | YOLO26x-sem | YOLO26l-sem | CABiNet (MobileNetV3-Large) | YOLO26m-sem | YOLO26s-sem | YOLO26n-sem | | --- | --- | --- | --- | --- | --- | --- | | Background | 81.69 | 79.43 | 82.67 | 80.75 | 79.31 | 78.43 | | Person | 53.34 | 49.28 | 46.9 | 50.81 | 49.44 | 44.51 | | Bike | 35.61 | 42.18 | 42.35 | 38.4 | 30.09 | 35.7 | | Car | 86.16 | 87.67 | 85.08 | 84.76 | 84.54 | 85.28 | | Drone | 60.46 | 65.97 | 61.77 | 58.91 | 59.73 | 50.88 | | Boat | 82.04 | 79.76 | 77.67 | 80.33 | 76.19 | 80.28 | | Animal | 41.21 | 37.08 | 33.74 | 38.22 | 36.21 | 37.15 | | Obstacle | 26.91 | 29.17 | 23.13 | 24.14 | 19.8 | 17.29 | | Construction | 75.76 | 67.09 | 78.81 | 72.87 | 71.02 | 74.39 | | Vegetation | 94.07 | 93.6 | 94.3 | 93.99 | 93.93 | 93.55 | | Road | 92.14 | 92.07 | 92.91 | 91.94 | 91.64 | 89.1 | | Sky | 90.96 | 92.72 | 94.57 | 88.57 | 93.35 | 91.73 | --- ## Evaluation Visualizations ### Per-Class IoU Bar Chart ![IoU Bar Chart](iou_bar_chart.png) ### Confusion Matrix ![Confusion Matrix](confusion_matrix_normalized.png) ### Loss Curves ![Loss Curves](results.png) --- ## Dataset [AeroScapes](https://github.com/ishann/aeroscapes) is an aerial semantic segmentation benchmark of ~3,200 720p images captured by a fleet of drones across roughly 10 flight sequences, covering urban and suburban scenes with people, vehicles, and small objects. ### Classes - Background - Person - Bike - Car - Drone - Boat - Animal - Obstacle - Construction - Vegetation - Road - Sky --- ## Usage ### Install Dependencies ```bash pip install ultralytics huggingface_hub ``` ### Load Model from Hugging Face ```python from huggingface_hub import hf_hub_download from ultralytics import YOLO weights = hf_hub_download( repo_id="dronefreak/aeroscapes-yolo26s-sem", filename="best.pt" ) model = YOLO(weights) ``` ### Run Inference ```python results = model.predict(source="image.png", task="semantic", imgsz=896) mask = results[0].semantic_mask.cpu().numpy().data # (H, W) class-ID map ``` --- ## Training Configuration | Setting | Value | | ------------ | ------------------------------------------ | | Epochs | 120 | | Image size | 896 | | Batch size | 8 | | Dataset | AeroScapes (converted images/+masks/ format) | | Framework | Ultralytics YOLO | | cls_pw (class weighting) | 0.5 | --- ## Official Resources - **AeroScapes Semantic Segmentation Model Zoo:** https://huggingface.co/collections/dronefreak/aeroscapes-semantic-segmentation-model-zoo - **CABiNet repository:** https://github.com/dronefreak/CABiNet - **CABiNet Paper:** https://arxiv.org/abs/2011.00993v2 - **AeroScapes Dataset (Hugging Face):** https://huggingface.co/datasets/dronefreak/Aeroscapes - **AeroScapes Repository:** https://github.com/ishann/aeroscapes - **AeroScapes Paper (CMU PDF):** https://www.cs.cmu.edu/~deva/papers/aeroscapes.pdf - **Ultralytics YOLO:** https://github.com/ultralytics/ultralytics - **Ultralytics YOLO26 Paper:** https://arxiv.org/abs/2606.03748 --- ## Training Framework Trained with the [CABiNet repository](https://github.com/dronefreak/CABiNet), which pairs its own real-time segmentation trainer with a parallel Ultralytics YOLO26-sem pipeline — shared dataset tooling, training/eval, and mIoU benchmarking across UAVid, AeroScapes, and VDD. Star the repo if you find these models useful! --- ## Known Limitations Performance may degrade in: * Very rare classes (Bike, Drone, Boat) each cover well under 2% of any single frame in the val split — expect unreliable predictions for these * No official test split — all reported numbers are on the val split, unlike UAVid/VDD * Frames are drawn from only ~10 flight sequences, so scene diversity is lower than UAVid/VDD despite the image count * Native resolution (1280x720) downsampled to the eval imgsz — fine detail on small objects (pedestrians, bikes) can be lost --- ## Citation Please cite the following: ```bibtex @inproceedings{nigam2018ensemble, title={Ensemble Knowledge Transfer for Semantic Segmentation}, author={Nigam, Ishan and Huang, Chen and Ramanan, Deva}, booktitle={2018 IEEE Winter Conference on Applications of Computer Vision (WACV)}, year={2018} } @INPROCEEDINGS{9560977, author={Kumaar, Saumya and Lyu, Ye and Nex, Francesco and Yang, Michael Ying}, booktitle={2021 IEEE International Conference on Robotics and Automation (ICRA)}, title={CABiNet: Efficient Context Aggregation Network for Low-Latency Semantic Segmentation}, year={2021}, pages={13517-13524}, doi={10.1109/ICRA48506.2021.9560977} } @article{Kumaar_Real-time_Semantic_Segmentation_2021, author = {Kumaar, Saumya and Lyu, Ye and Nex, Francesco and Yang, Michael Ying}, doi = {10.1016/j.isprsjprs.2021.06.006}, journal = {ISPRS Journal of Photogrammetry and Remote Sensing}, pages = {124--134}, title = {{Real-time Semantic Segmentation with Context Aggregation Network}}, url = {https://www.sciencedirect.com/science/article/pii/S0924271621001647}, volume = {178}, year = {2021} } @article{jocher2026ultralytics, title={Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models}, author={Jocher, Glenn and Qiu, Jing and Liu, Mengyu and Lyu, Shuai and Akyon, Fatih Cagatay and Kalfaoglu, Muhammet Esat}, journal={arXiv preprint arXiv:2606.03748}, year={2026} } @software{cabinet_uavid_benchmark, author = {Kumaar, Saumya}, title = {CABiNet: Semantic Segmentation Benchmarking on UAVid (CABiNet vs. YOLO26)}, url = {https://github.com/dronefreak/CABiNet}, year = {2026} } ```