基于YOLOv3的正下无人机视角挖掘机实时检测方法
PDF下载 (3338)蔡振宇,王泽锴,陈特欢,李文来.基于YOLOv3的正下无人机视角挖掘机实时检测方法[J].宁波大学学报(理工版),2021,34(2):42-48.DOI:
CAI Zhenyu,WANG Zekai,CHEN Tehuan,LI Wenlai.Real-time excavator detection under direct UAV view based on improved YOLOv3 method[J].Journal of Ningbo University(Natural Science & Engineering Edition),2021,34(2):42-48.DOI:
| Title: | Real-time excavator detection under direct UAV view based on improved YOLOv3 method |
| 作者: | 蔡振宇, 王泽锴, 陈特欢, 李文来 |
| Author(s): | CAI Zhenyu, WANG Zekai, CHEN Tehuan, LI Wenlai |
| 关键词: | 无人机视角; 特征融合; 目标检测 |
| Keywords: | YOLOv3; UAV view; feature fusion; object detection |
| 分类号: | TP391.4 |
| 文献标识码: | A |
| 摘要: | 针对无人机巡检的智能化要求, 提出一种应对高空巡检场景下的实时挖掘机检测模型. 该模型以YOLOv3为基础, 将骨干网络精简至43层, 通过特征融合策略使检测任务在两个尺度上进行. 此外模型还借鉴了focal loss的思想设计损失函数. 文中实验对象为正下无人机视角的挖掘机目标. 在完成了数据集的搭建工作后, 根据正下无人机视角的目标特性进行训练, 使模型达到最优解. 最终经实验验证, 在相同输入尺寸的情况下, 本文所提出的检测模型比YOLOv3准确率更高、鲁棒性更好, 且帧数可提升10帧●s-1. |
| Abstract: | A real-time excavator detection method for high-altitude patrol scenarios is proposed to meet the requirements of Unmanned Aerial Vehicle (UAV) intelligent inspection. The model is based on YOLOv3 and adopts a 43-layer backbone network. The object can be detected on two scales through feature fusion strategy. Besides, it also defines the loss function with the idea of focal loss. The experimental object in this article is an excavator under the direct UAV view. Upon the completion of the construction of the dataset, the optimal solution of the model is obtained by training the model with characteristics of the object under direct UAV view. The experimental results show that the proposed model is more accurate and robust than YOLOv3 for the same input size, and has increased by 10 frame per second in terms of frame number. |
| 参考文献 /References: | [1] 陈志强. 浅谈变电站机器人巡检系统优化应用[J]. 工程技术(文摘版)建筑, 2016(2):156. [2] Ma X X, Grimson W E L. Edge-based rich representation for vehicle classification[C]//Tenth IEEE International Conference on Computer Vision, 2005:1185-1192. [3] Dalal N, Triggs B. Histograms of oriented gradients for human detection[C]//2005 IEEE Conference on Computer Vision and Pattern Recognition, 2005:886-893. [4] Kazemi F M, Samadi S, Poorreza H R, et al. Vehicle recognition using curvelet transform and SVM[C]//Fourth International Conference on Information Technology, 2007:516-521. [5] Freund Y, Schapire R E. A decision-theoretic generalization of on-line learning and an application to boosting[J]. Journal of Computer and System Sciences, 1997, 55(1):119-139. [6] 王健, 王晓东, 郭磊. 基于混合高斯模型的运动目标检测[J]. 宁波大学学报(理工版), 2019, 32(3):35-39. [7] Hinton G E, Salakhutdinov R. Reducing the dimensionality of data with neural networks[J]. Science, 2006, 313(5786):504-507. [8] 周晓杰, 蔡元强, 夏克江, 等. 基于火焰图像显著区域特征学习与分类器融合的回转窑烧结工况识别[J]. 控制与决策, 2017, 32(1):187-192. [9] Liu W, Anguelov D, Erhan D, et al. SSD: Single shot multibox detector[C]//European Conference on Computer Vision, 2016:21-37. [10] Redmon J, Divvala S, Girshick R, et al. You only look once: Unified, real-time object detection[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition, 2016:779-788. [11] Redmon J, Farhadi A. YOLO9000: Better, faster, stronger [C]//2017 IEEE Conference on Computer Vision and Pattern Recognition, 2017:6517-6525. [12] Redmon J, Farhadi A. YOLOv3: An incremental improvement[EB/OL]. [2019-08-17]. https://arxiv.org/abs/ 1804.02767. [13] Lin T, Goyal P, Girshick R, et al. Focal loss for dense object detection[C]//International Conference on Computer Vision, 2017:2999-3007. [14] 张富凯, 杨峰, 李策. 基于改进YOLOv3的快速车辆检测方法[J]. 计算机工程与应用, 2019, 55(2):12-20. [15] 戴伟聪, 金龙旭, 李国宁, 等. 遥感图像中飞机的改进YOLOv3实时检测算法[J]. 光电工程, 2018, 45(12):84- 92. [16] 舒军, 吴柯. 基于改进YOLOv3的航拍目标实时检测方法[J]. 湖北工业大学学报, 2020(1):21-24. [17] He K, Zhang X, Ren S, et al. Deep residual learning for image recognition[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition, 2016:770-778. [18] Lin T, Dollar P, Girshick R, et al. Feature pyramid networks for object detection[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition, 2017:936- 944. [19] Ioffe S, Szegedy C. Batch normalization: Accelerating deep network training by reducing internal covariate shift [EB/OL]. [2019-07-12]. https://arxiv.org/abs/1502.03167. [20] Everingham M, Eslami S M A, van Gool L, et al. The pascal visual object classes challenge: A retrospective[J]. International Journal of Computer Vision, 2015, 111(1):98-136. [21] Ren S, He K, Girshick R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6):1137-1149. |
| 备注/Memo: | 收稿日期: 2020-06-06. 宁波大学学报(理工版)网址: http://journallg.nbu.edu.cn/ 基金项目: 国家自然科学基金(61703217); 浙江省教育厅科研项目(Y201840096, Y201839158, Y201840072); 宁波市科技创新2025重大专项(2019B10100). 第一作者: 蔡振宇(1994-), 男, 湖南株洲人, 在读硕士研究生, 主要研究方向: 深度学习、目标检测. E-mail: bnbnnvh@163.com *通信作者: 陈特欢(1988-), 男, 浙江宁波人, 副教授, 主要研究方向: 微流体控制、计算最优控制、机器人控制. E-mail: chentehuan@nbu.edu.cn 宁波大学学报(理工版)网址:http://journallg.nbu.edu.cn/ |