Submitted:
28 October 2025
Posted:
29 October 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We propose an Adaptive Attention Fusion (AAF) module that adaptively integrates sparse and dense attention through a pre-fusion mechanism. This design enhances the model's ability to extract global contextual features, thereby significantly improving the detection performance on small objects.
- To further enhance the detection capability for small objects, we introduce a novel feature fusion pyramid network, termed EMS-FPN. This module enlarges the receptive field and strengthens detail preservation through globally heterogeneous multi-scale convolution, all while maintaining low model complexity in terms of parameters and computational cost.
- We design a Lightweight Attention-Gated Module (LAGM) based on gated inverted bottleneck convolution to achieve locally dynamic perception. LAGM improves feature modeling in dense scenes and for small objects, while preserving the overall efficiency and lightweight nature of the model.
2. Related Work
3. Methods
3.1. Overview
3.2. Adaptive Attention Fusion (AAF)
3.3. Efficient Multi-Scale Semantic FPN (EMS-FPN)
3.4. Local-Aware Gated Module (LAGM)
4. Experiments
4.1. Datasets and Experimental Setup
4.2. Ablation Experiment
4.3. Comparisons with Other Object Detection Networks
4.4. Visualization
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| DETR | Detection Transformer |
| AAF | Adaptive Attention Fusion |
| FPN | Feature Pyramid Networks |
| LAGM | Local-Aware Gated Module |
References
- Feng, J.; Wang, J.; Qin, R. Lightweight detection network for arbitrary oriented vehicles in UAV imagery via precise positional information encoding and bidirectional feature fusion. International Journal of Remote Sensing, 2023, *44*, 4529–4558. [Google Scholar]
- Bhadra, S.; Sagan, V.; Sarkar, S.; Braud, M.; Mockler, T.C.; Eveland, A.L. Prosail-net: A transfer learning-based dual stream neural network to estimate leaf chlorophyll and leaf angle of crops from UAV hyperspectral images. ISPRS Journal of Photogrammetry and Remote Sensing 2024, *210*, 1–24.
- Alexan, W.; Aly, L.; Korayem, Y.; Gabr, M.; El-Damak, D.; Fathy, A.; Mansour, H.A.A. Secure communication of military reconnaissance images over UAV-assisted relay networks. IEEE Access 2024, *12*, 78589–78610.
- Cao, Z.; Kooistra, L.; Wang, W.; Guo, L.; Valente, J. Real-time object detection based on UAV remote sensing: A systematic literature review. Drones 2023, *7*, 620. Available online: https://www.mdpi.com/2504-446X/7/10/620.
- Zhang, Z. Drone-yolo: An efficient neural network method for target detection in drone images. Drones 2023, *7*, 526. Available online: https://www.mdpi.com/2504-446X/7/8/526.
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 580–587. [Google Scholar]
- Girshick, R. Fast R-CNN. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 7–13 December 2015; pp. 1440–1448. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2016, *39*, 1137–1149. [Google Scholar]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar]
- Redmon, J.; Farhadi, A. YOLOv3: An incremental improvement. arXiv 2018, arXiv:1804.02767. [Google Scholar] [CrossRef]
- Bochkovskiy, A.; Wang, C.-Y.; Liao, H.-Y.M. YOLOv4: Optimal speed and accuracy of object detection. arXiv 2020, arXiv:2004.10934. [Google Scholar] [CrossRef]
- Jocher, G.; Chaurasia, A.; Stoken, A.; Borovec, J.; Kwon, Y.; Michael, K.; Fang, J.; Wong, C.; Yifu, Z.; Montes, D.; et al. ultralytics/yolov5: v6.2 - YOLOv5 classification models, Apple M1, reproducibility, ClearML and Deci.ai integrations. Zenodo, 2022. [Google Scholar]
- Li, C.; Li, L.; Jiang, H.; Weng, K.; Geng, Y.; Li, L.; Ke, Z.; Li, Q.; Cheng, M.; Nie, W.; et al. YOLOv6: A single-stage object detection framework for industrial applications. arXiv 2022, arXiv:2209.02976. [Google Scholar] [CrossRef]
- Wang, C.-Y.; Bochkovskiy, A.; Liao, H.-Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 7464–7475. [Google Scholar]
- Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLO (Version 8.0.0) [Computer Software]. 2023.
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In European Conference on Computer Vision; Springer: Glasgow, UK, 23-28 August 2020; pp. 213–229. [Google Scholar]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable DETR: Deformable transformers for end-to-end object detection. arXiv 2020, arXiv:2010.04159. [Google Scholar]
- Li, F.; Zhang, H.; Liu, S.; Guo, J.; Ni, L.M.; Zhang, L. DN-DETR: Accelerate DETR training by introducing query denoising. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 13619–13627. [Google Scholar]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs beat YOLOs on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 16965–16974. [Google Scholar]
- Zhai, X.; Huang, Z.; Li, T.; Liu, H.; Wang, S. YOLO-Drone: An optimized YOLOv8 network for tiny UAV object detection. Electronics, 2003, *12*, 3664. [Google Scholar]
- Sun, W.; Dai, L.; Zhang, X.; Chang, P.; He, X. RSOD: Real-time small object detection algorithm in UAV-based traffic monitoring. Applied Intelligence, 2022, *53*, 1–16. [Google Scholar]
- Wang, G.; Chen, Y.; An, P.; Hong, H.; Hu, J.; Huang, T. UAV-YOLOv8: A small-object-detection model based on improved YOLOv8 for UAV aerial photography scenarios. Sensors, 2023, *23*, 7190. [Google Scholar]
- Zeng, S.; Yang, W.; Jiao, Y.; Geng, L.; Chen, X. SCA-YOLO: A new small object detection model for UAV images. The Visual Computer, 2024, *40*, 1787–1803. [Google Scholar]
- Zhao, H.; Zhang, H.; Zhao, Y. YOLOv7-Sea: Object detection of maritime UAV images based on improved YOLOv7. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 2–7 January 2023; pp. 233–238. [Google Scholar]
- Zhang, H.; Liu, K.; Gan, Z.; Zhu, G.-N. UAV-DETR: Efficient end-to-end object detection for unmanned aerial vehicle imagery. arXiv 2025, arXiv:2501.01855. [Google Scholar]
- Kong, Y.; Shang, X.; Jia, S. Drone-DETR: Efficient small object detection for remote sensing image using enhanced RT-DETR model. Sensors, 2024, *24*, 5496. [Google Scholar]
- Liu, Y.; He, M.; Hui, B. ESO-DETR: An improved real-time detection transformer model for enhanced small object detection in UAV imagery. Drones, 2025, *9*, 143. [Google Scholar]
- Zhou, S.; Chen, D.; Pan, J.; Shi, J.; Yang, J. Adapt or perish: Adaptive sparse transformer with attentive feature refinement for image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 2952–2963. [Google Scholar]
- Tan, M.; Pang, R.; Le, Q.V. EfficientDet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 10781–10790. [Google Scholar]
- Rahman, M.M.; Munir, M.; Marculescu, R. EMCAD: Efficient multi-scale convolutional attention decoding for medical image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 11769–11779. [Google Scholar]
- Li, Y.; Chen, Y.; Wang, N.; Zhang, Z. Scale-aware trident networks for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 6054–6063. [Google Scholar]
- Song, Y.; Zhou, Y.; Qian, H.; Du, X. Rethinking performance gains in image dehazing networks. arXiv 2022, arXiv:2209.11448. [Google Scholar] [CrossRef]
- Chollet, F. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1251–1258. [Google Scholar]
- Du, D.; Zhu, P.; Wen, L.; Bian, X.; Lin, H.; Hu, Q.; Peng, T.; Zheng, J.; Wang, X.; Zhang, Y.; et al. VisDrone-DET2019: The vision meets drone object detection in image challenge results. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, Seoul, Republic of Korea, 27–28 October 2019; pp. 0–0. [Google Scholar]
- Tang, S.; Zhang, S.; Fang, Y. HIC-YOLOv5: Improved YOLOv5 for small object detection. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation, Yokohama, Japan, 13–17 May 2024; pp. 6614–6619. [Google Scholar]
- Liu, X.; Zhou, S.; Ma, J.; Sun, Y.; Zhang, J.; Zuo, H. DFAS-YOLO: Dual Feature-Aware Sampling for Small-Object Detection in Remote Sensing Images. Remote Sens. 2025, 17, 3476. [Google Scholar] [CrossRef]
- Roh, B.; Shin, J.; Shin, W.; Kim, S. Sparse DETR: Efficient end-to-end object detection with learnable sparsity. arXiv 2021, arXiv:2111.14330. [Google Scholar] [CrossRef]








| Environment | Specification |
|---|---|
| CPU | Intel(R) Xeon(R) Platinum 8352 |
| GPU | vGPU-32GB |
| VRAM | 32GB |
| RAM | 90GB |
| Operating System | Ubuntu 22.04 |
| Language | Python 3.12 |
| Framework | PyTorch 2.5.1 |
| CUDA Version | 12.4 |
| Environment | Specification |
|---|---|
| Input Size | 640×640 |
| Batch Size | 16 |
| Training Epochs | 300 |
| Optimizer | AdamW |
| Initial Learning Rate | 0.0001 |
| Learning Rate Factor | 0.01 |
| Momentum | 0.9 |
| Warmup Steps | 2000 |
| Data | PG-Backbone | AAF | EMS-FPN | P | R | mAP@50 | Params | FLOPs |
|---|---|---|---|---|---|---|---|---|
| Val | – | – | – | 59.1 | 43.8 | 45.0 | 20.1M | 58.7G |
| √ | – | – | 57.9 | 42.9 | 44.1 | 10.0M | 30.7G | |
| √ | √ | – | 59.2 | 43.1 | 45.1 | 10.8M | 31.6G | |
| √ | √ | √ | 61.0 | 45.2 | 46.3 | 9.2M | 27.3G | |
| Test | – | – | – | 52.0 | 36.3 | 34.5 | 20.1M | 58.7G |
| √ | – | – | 52.0 | 35.4 | 33.7 | 10.0M | 30.7G | |
| √ | √ | – | 53.5 | 36.7 | 35.0 | 10.8M | 31.6G | |
| √ | √ | √ | 54.1 | 37.3 | 36.5 | 9.2M | 27.3G |
| Data | PG-Backbone | AAF | EMS-FPN | P | R | mAP@50 | Params | FLOPs |
|---|---|---|---|---|---|---|---|---|
| Val | – | – | – | 60.9 | 46.5 | 47.9 | 42.8M | 134.5G |
| √ | – | – | 60.2 | 46.0 | 47.2 | 13.7M | 50.0G | |
| √ | √ | – | 60.9 | 46.8 | 48.1 | 14.6M | 50.9G | |
| √ | √ | √ | 61.7 | 47.9 | 49.2 | 14.7M | 57.0G | |
| Test | – | – | – | 54.2 | 38.3 | 37.5 | 42.8M | 134.5G |
| √ | – | – | 53.7 | 37.8 | 36.7 | 13.7M | 50.0G | |
| √ | √ | – | 55.4 | 38.8 | 38.1 | 14.6M | 50.9G | |
| √ | √ | √ | 56.8 | 40.4 | 39.2 | 14.7M | 57.0G |
| Model | Input Size | mAP@50 | mAP@50:95 | Params | FLOPs |
|---|---|---|---|---|---|
| YOLOv8-M | 640×640 | 40.7 | 24.6 | 25.9M | 78.9G |
| YOLOv8-L | 640×640 | 42.7 | 26.1 | 43.7M 1 | 165.2G |
| YOLOv11-S | 640×640 | 38.7 | 23.0 | 9.4M | 21.3G |
| YOLOv11-M | 640×640 | 43.1 | 25.9 | 20.0M | 67.7G |
| Drone-YOLO-S | 640×640 | 44.3 | 27.0 | 10.9M | – |
| HIC-YOLOv5 | 640×640 | 44.3 | 26.0 | 9.4M | 31.2G |
| DFAS-YOLO | 640×640 | 44.8 | 27.3 | 7.5M | – |
| Deformable DETR | 1333×800 | 43.1 | 27.1 | 40.0M | 173.0G |
| RT-DETR-R18 | 640×640 | 45.0 | 27.2 | 20.1M | 58.7G |
| RT-DETR-R50 | 640×640 | 47.9 | 29.3 | 42.8M | 134.5G |
| UAV-DETR-R18 | 640×640 | 48.8 | 29.8 | 20.0M | 77.0G |
| LEA-DETR-S(Ours) | 640×640 | 46.3 | 28.5 | 9.2M | 27.3G |
| LEA-DETR-M(Ours) | 640×640 | 49.2 | 30.5 | 14.7M | 57.0G |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).