基于改进3D卷积网络的人体动作识别
PDF下载 (315)高海玲,王晓东,章联军,赵伸豪,金建国.基于改进3D卷积网络的人体动作识别[J].宁波大学学报(理工版),2023,36(3):16-21.DOI:
GAO Hailing,WANG Xiaodong,ZHANG Lianjun,ZHAO Shenhao,JIN Jianguo.Human motion recognition based on improved 3D convolution network[J].Journal of Ningbo University(Natural Science & Engineering Edition),2023,36(3):16-21.DOI:
| Title: | Human motion recognition based on improved 3D convolution network |
| 作者: | 高海玲, 王晓东, 章联军, 赵伸豪, 金建国 |
| Author(s): | GAO Hailing, WANG Xiaodong, ZHANG Lianjun, ZHAO Shenhao, JIN Jianguo |
| 关键词: | 深度学习; 人体动作识别; 3D卷积; 注意力机制 |
| Keywords: | deep learning; human motion recognition; 3D convolution; attention mechanism |
| 分类号: | TP391.4 |
| 文献标识码: | A |
| 摘要: | 为解决现有多数视频人体动作识别3D卷积方法无法区分信息中各维度的重要和非重要特征问题, 提出了通过门控循环单元(Gated Recurrent Unit, GRU)和空间注意力增强模块构建时空特征处理网络的方法, 基于多级特征融合和多组通道注意力特征选择构建网络, 改进基础网络模型ResNet3D对视频人体动作识别中的网络模型. 改进后模型在2个公开数据集UCF101和HMDB51上的准确率分别为96.42%和71.08%, 与C3D、Two-stream等网络模型相比, 具有更高的识别准确率. |
| Abstract: | Video human motion recognition research has great potential for applications, but the modeling quality is greatly affected by movement types, environmental differences and other factors. Most 3D convolution methods for video human motion recognition cannot distinguish between important and non-important features in each dimension given the needed information. To tackle this problem the GRU gating unit and spatial attention enhancement module are used to build a spatio-temporal feature processing network, and the network is built based on multi-level feature fusion and multi-channel attention feature selection. Based on the basic network model ResNet3D, the network model in video human motion recognition is improved. The model achieves 96.42% and 71.08% recognition accuracy on two public datasets UCF101 and HMDB51, respectively, with satisfactory recognition performance. Compared with C3D, two-stream and other generic network models, the proposed model shows higher recognition accuracy, which indicates the effectiveness of the proposed model. |
| 参考文献 /References: | [1] Wang H, Schmid C. Action recognition with improved trajectories[C]//2013 IEEE International Conference on Computer Vision, Sydney, Australia, 2014:3551-3558. [2] Simonyan K, Zisserman A. Two-stream convolutional networks for action recognition in videos[EB/OL]. [2022-08-14]. https://arxiv.org/abs/1406.2199. [3] Ji S W, Xu W, Yang M, et al. 3D convolutional neural networks for human action recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2013, 35(1):221-231. [4] Tran D, Bourdev L, Fergus R, et al. Learning spatiotemporal features with 3D convolutional networks [C]//2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 2016:4489-4497. [5] Diba A L, Fayyaz M, Sharma V, et al. Temporal 3D ConvNets: New architecture and transfer learning for video classification[EB/OL]. [2022-08-14]. https://arxiv.org/abs/ 1711.08200. [6] Hara K, Kataoka H, Satoh Y. Can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet? [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 6546-6555. [7] Carreira J, Zisserman A. Quo vadis, action recognition? A new model and the kinetics dataset[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, USA, 2017:4724-4733. [8] 张瑞, 李其申, 储. 基于3D卷积神经网络的人体动作识别算法[J]. 计算机工程, 2019, 45(1):259-263. [9] 刘悦, 张雷, 辛山, 等. 融入时空注意力机制的深度学习网络视频动作分类[J]. 中国科技论文, 2022, 17(3): 281-287. [10] Zolfaghari M, Singh K, Brox T. ECO: Efficient convolutional network for online video understanding [C]//European Conference on Computer Vision, Cham: Springer, 2018:713-730. [11] Feichtenhofer C, Fan H Q, Malik J, et al. SlowFast networks for video recognition[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South), 2020:6201-6210. [12] Li Y, Ji B, Shi X T, et al. TEA: Temporal excitation and aggregation for action recognition[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2020:906-915. [13] Wang L M, Tong Z, Ji B, et al. TDN: Temporal difference networks for efficient action recognition[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, USA, 2021:1895-1904. [14] Yi Z W, Sun Z H, Feng J C, et al. 3D residual networks with channel-spatial attention module for action recognition[C]//2020 Chinese Automation Congress (CAC), Shanghai, China, 2021:5171-5174. [15] 张聪聪, 何宁, 孙琪翔, 等. 基于注意力机制的3D DenseNet人体动作识别方法[J]. 计算机工程, 2021, 47(11):313-320. [16] Zhu L C, Fan H H, Luo Y W, et al. Temporal cross-layer correlation mining for action recognition[J]. IEEE Transactions on Multimedia, 2022, 24:668-676. [17] Chung J, Gulcehre C, Cho K, et al. Empirical evaluation of gated recurrent neural networks on sequence modeling [EB/OL]. [2022-08-14]. https://arxiv.org/abs/1412.3555. [18] Szegedy C, Vanhoucke V, Ioffe S, et al. Rethinking the inception architecture for computer vision[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, 2016:2818-2826. [19] Soomro K, Amir R Z, Shah M. UCF101: A dataset of 101 human actions classes from videos in the wild[EB/OL]. [2022-08-14]. http://export.arxiv.org/pdf/1212.0402. [20] Kuehne H, Jhuang H, Garrote E, et al. HMDB: A large video database for human motion recognition[C]// Proceedings of 2011 IEEE International Conference on Computer Vision, USA: IEEE Press, 2011:2556-2563. |
| 备注/Memo: | 收稿日期:2022-10-14.宁波大学学报(理工版)网址:http://journallg.nbu.edu.cn/ 基金项目:浙江省自然科学基金(LY20F010005);宁波市“科技创新2025”重大专项(2022T005). 第一作者:高海玲(1998-),女,甘肃白银人,在读硕士研究生,主要研究方向:视频信息处理.E-mail:1450363642@qq.com *通信作者:王晓东(1970-),男,浙江上虞人,教授,主要研究方向:多媒体信号处理.E-mail:wxd@nbu.edu.cn 宁波大学学报(理工版)网址:http://journallg.nbu.edu.cn/ |