RepViT-Based UAV Small Object Detection on VisDrone and UAVDT
view PDF
view PDF

How to Cite

M., Jiva Bharathi, and Sathya M. 2026. “RepViT-Based UAV Small Object Detection on VisDrone and UAVDT”. Journal of Electronics and Informatics 8 (3): 246-60. https://doi.org/10.36548/jei.2026.3.004.

Keywords

UAV Object Detection
RepViT
Feature Pyramid Network (FPN)
VisDrone
UAVDT (Unmanned Aerial Vehicle Detection and Tracking)

Abstract

The Unmanned Aerial Vehicles (UAVs), popularly referred to as drones, have gained increasing significance with respect to their use in applications like traffic monitoring, surveillance, disaster management, environmental observation, and precision agriculture. Despite substantial progress in aerial image sensing technologies, accurate object detection in UAV imagery has remained elusive owing to the issues related to small object sizes, object scale variance, motion blur, occlusions, and complexity of the background. The proposed study describes a RepViT architecture to enhance the detection of small objects. The proposed framework makes use of RepViT architecture, Feature Pyramid Network (FPN), and attention mechanism for enhancing discriminative features and suppressing background noise. The effectiveness of the proposed model was evaluated using benchmark UAV object detection datasets such as VisDrone and UAVDT using widely adopted object detection metrics. The proposed framework has sustained the inference rate of 48 frames per second (FPS).

References

  1. Redmon, Joseph, Santosh Divvala, Ross Girshick, and Ali Farhadi. "You Only Look Once: Unified, Real-Time Object Detection." In Proceedings of the IEEE conference on computer vision and pattern recognition 2016, 779-788.
  2. Redmon, Joseph, and Ali Farhadi. "YOLO9000: Better, Faster, Stronger." In Proceedings of the IEEE conference on computer vision and pattern recognition 2017, 7263-7271.
  3. Redmon, Joseph. "Yolov3: An Incremental Improvement." arXiv preprint arXiv:1804.02767 (2018).
  4. Bochkovskiy, Alexey, Chien-Yao Wang, and Hong-Yuan Mark Liao. "Yolov4: Optimal Speed And Accuracy of Object Detection." arXiv preprint arXiv:2004.10934 (2020).
  5. Wang, Chien-Yao, Alexey Bochkovskiy, and Hong-Yuan Mark Liao. "YOLOv7: Trainable Bag-Of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors." In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition 2023, 7464-7475.
  6. Jocher, Glenn, Liu Changyu, Adam Hogan, Lijun Yu, Prashant Rai, and Trevor Sullivan. "ultralytics/yolov5: Initial Release." Zenodo 2020.
  7. He, Kaiming, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. "Mask R-Cnn." In Proceedings of the IEEE international conference on computer vision 2017, 2961-2969.
  8. Ren, Shaoqing, Kaiming He, Ross Girshick, and Jian Sun. "Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks." IEEE transactions on pattern analysis and machine intelligence 2016, vol. 39, no. 6, 1137-1149.
  9. Lin, Tsung-Yi, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. "Feature Pyramid Networks for Object Detection." In Proceedings of the IEEE conference on computer vision and pattern recognition 2017, 2117-2125.
  10. Ashish, Vaswani. "Attention is All You Need." Advances in neural information processing systems 2017, vol. 30, I.
  11. Dosovitskiy, Alexey, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani "An Image is Worth 16x16 Words: Transformers For Image Recognition at Scale." arXiv preprint arXiv:2010.11929 (2020).
  12. Liu, Ze, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. "Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows." In Proceedings of the IEEE/CVF international conference on computer vision 2021, 10012-10022.
  13. Kaiming, He, Zhang Xiangyu, Ren Shaoqing, and Sun Jian. "Deep Residual Learning for Image Recognition." In Proceedings of the IEEE conference on computer vision and pattern recognition 2016, vol. 34, 770-778.
  14. Wang, Ao, Hui Chen, Zijia Lin, Jungong Han, and Guiguang Ding. "Repvit: Revisiting Mobile Cnn From Vit Perspective." In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition 2024, 15909-15920.
  15. Sandler, Mark, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. "Mobilenetv2: Inverted Residuals and Linear Bottlenecks." In Proceedings of the IEEE conference on computer vision and pattern recognition 2018, 4510-4520.
  16. VisDrone Dataset - https://www.kaggle.com/datasets/banuprasadb/visdrone-dataset
  17. UAVDT Dataset - https://www.kaggle.com/datasets/shakaibkaggle/uavdt-dataset