Submitted:
15 October 2024
Posted:
15 October 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Skeleton Behavior Recognition Based on Graph Convolution
2.1. Spatio-Temporal Graph Convolutional Network
2.2. Analysis of Bias Weighting Problem Methods
3. Improved Graph Convolutional Human Behavior Recognition Algorithm
3.1. The Multi-Scale Spatio-Temporal Graph Convolution Network Incorporating Multi-Granularity Features
3.2. Skeleton Fine-Grained Partitioning Strategy
3.3. Cross-Scale Feature Fusion Layer
4. Experiment and Result Analysis
4.1. Experimental Dataset
4.2. Experimental Environment and Settings
4.3. Experimental Results and Analysis
4.3.1. Comparative Experiment Using Unbiased Weighting Method
4.3.2. Comparative Experiments on Fusing Multi-Granularity Features
4.3.4. Comparison Experiment with other Models
5. Conclusion
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Pilarski, P.M.; Butcher, A.; Johanson, M.; Botvinick, M.M.; Bolt, A.; Parker, A.S. Learned human-agent decision-making, communication and joint action in a virtual reality environment. arXiv 2019, arXiv:1905.02691. [Google Scholar]
- Shi, L.; Zhou, Y.; Wang, J.; et al. Compact global association based adaptive routing framework for personnel behavior understanding. Future Generation Computer Systems 2023, 141, 514–525. [Google Scholar] [CrossRef]
- Sudha, M.R.; Sriraghav, K.; Jacob, S.G.; et al. Approaches and applications of virtual reality and gesture recognition: A review. International Journal of Ambient Computing and Intelligence (IJACI) 2017, 8, 1–18. [Google Scholar] [CrossRef]
- Duan, H.; Zhao, Y.; Chen, K.; et al. Revisiting skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2022; pp. 2969–2978. [Google Scholar]
- Liu, C.; Ying, J.; Yang, H.; Hu, X.; Liu, J. Improved human action recognition approach based on two-stream convolutional neural network model. The Visual Computer 2021, 37, 1327–1341. [Google Scholar] [CrossRef]
- Duan, H.; Zhao, Y.; Chen, K. Revisiting Skeleton-Based Action Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2022; pp. 2969–2978. [Google Scholar]
- Liu, J.; Shahroudy, A.; Xu, D.; et al. Spatio-temporal LSTM with trust gates for 3D human action recognition. In Proceedings of the 14th European Conference on Computer Vision; Springer: Heidelberg; 2016; pp. 816–833. [Google Scholar]
- Wei, S.H.; Song, Y.H.; Zhang, Y.L. Human skeleton tree recurrent neural network with joint relative motion feature for skeleton based action recognition. In Proceedings of the IEEE International Conference on Image Processing (ICIP); IEEE Computer Society Press: Los Alamitos; 2017; pp. 91–95. [Google Scholar]
- Zheng, W.; Li, L.; Zhang, Z.X.; et al. Relational network for skeleton-based action recognition. In Proceedings of the IEEE International Conference on Multimedia and Expo(ICME); IEEE Computer Society Press: Los Alamitos; 2019; pp. 826–831. [Google Scholar]
- Zhao, R.; Wang, K.; Su, H.; Ji, Q. Bayesian graph convolution lstm for skeleton-based action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019. [Google Scholar]
- Yan, S.; Xiong, Y.; Lin, D. Spatial temporal graph convolutional networks for skeleton-based action recognition. In Proceedings of the Thirty-second AAAI Conference on Artificial Intelligence; 2018. [Google Scholar]
- Shi, L.; Zhang, Y.; Cheng, J.; et al. Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, 2019; pp. 12026–12035. [Google Scholar]
- Li, M.S.; Chen, S.H.; Chen, X.; et al. Actional-structural graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE Computer Society Press: Los Alamitos; 2019; pp. 3590–3598. [Google Scholar]
- Li, C.; Cui, Z.; Zheng, W.; Xu, C.; Yang, J. Spatio-temporal graph convolution for skeleton based action recognition. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence; 2018. [Google Scholar]
- Li, M.; Chen, S.; Chen, X.; Zhang, Y.; Wang, Y.; Tian, Q. Actional-structural graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2019; pp. 3595–3603. [Google Scholar]
- Liao, R.; Zhao, Z.; Urtasun, R.; SZemel, R. Lanczosnet: Multi-scale deep graph convolutional networks. arXiv 2019, arXiv:1901.01484. [Google Scholar]
- Cheng, K.; Zhang, Y.; He, X.; et al. Skeleton-based action recognition with shift graph convolutional network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2020; pp. 183–192. [Google Scholar]
- Li, W.; Liu, X.; Liu, Z.; et al. Skeleton-based action recognition using multi-scale and multi-stream improved graph convolutional network. IEEE Access 2020, 8, 144529–144542. [Google Scholar] [CrossRef]
- Liu, Z.; Zhang, H.; Chen, Z.; et al. Disentangling and unifying graph convolutions for skeleton-based action recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition; IEEE: Piscataway, NJ, 2020; pp. 143–152. [Google Scholar]
- Yang, Y.; Deng, C.; Gao, S.; Liu, W.; Tao, D.; Gao, X. Discriminative multi-instance multitasks learning for 3D action recognition. IEEE Trans. Multimed. 2017, 19, 519–529. [Google Scholar] [CrossRef]
- Ran, X.Y.; Liu, K.; Li, G.; Ding, W.W.; Chen, B. Human action recognition algorithm based on adaptive skeleton center. Journal of Image and Graphics 2018, 23, 519–525. [Google Scholar]
- Agahian, S.; Negin, F.; Köse, C. Improving bag-of-poses with semi-temporal pose descriptors for skeleton-based action recognition. Vis. Comput. 2019, 35, 591–607. [Google Scholar] [CrossRef]
- Bulbul, M.F.; Tabussum, S.; Ali, H.; et al. Exploring 3D human action recognition using STACOG on multi-view depth motion graphs sequences. Sensors 2021, 21, 3642–3651. [Google Scholar] [CrossRef] [PubMed]
- Zhang, C.; Liang, J.; Li, X.; Xia, Y.; Di, L.; Hou, Z.; Huan, Z. Human action recognition based on enhanced data guidance and key node spatial temporal graph convolution. Multimed. Tools Appl. 2022, 81, 8349–8366. [Google Scholar] [CrossRef]
- Wu, Q.; Huang, Q.; Li, X. Multimodal human action recognition based on spatio-temporal action representation recognition model. Multimed. Tools Appl. 2022, 81, 1–22. [Google Scholar] [CrossRef]
- You, K.; Hou, Z.; Liang, J.; Lin, E.; Shi, H.; Zhong, Z. A 4D strong spatio-temporal feature learning network for behavior recognition of point cloud sequences. Multimedia Tools and Applications 2024. [CrossRef]






| Model methods | The value of K in K-order Adjacency Matrix | ||||||
|---|---|---|---|---|---|---|---|
| K = 2 | K = 3 | K = 4 | K = 5 | K = 6 | K = 8 | K = 10 | |
| MS-TGCN | 94.09 | 93.31 | 92.52 | 92.91 | 92.12 | 92.91 | 92.12 |
| MS-TGCN-D | 94.28 | 93.70 | 93.31 | 92.91 | 94.88 | 92.13 | 91.73 |
| Number of joint points / (pieces) | accuracy rate (%) | ||
|---|---|---|---|
| 20 | 23 | 25 | |
| ✓ | 94.88 | ||
| ✓ | 94.09 | ||
| ✓ | 93.31 | ||
| MSR Action 3D behavior types | Accuracy rate of behavior recognition (%) | |
|---|---|---|
| MS-TGCN-D (20-joints) | MS-TGCN-D (25-joints) | |
| Raise your hand high(HiW) | 100 | 81.8 |
| Wave your hand in front of your chest(HoW) | 100 | 100 |
| Hammering(H) | 92.3 | 75.0 |
| Hand catch(HCh) | 100 | 100 |
| Forward punch(FP) | 100 | 90.9 |
| High throw(HT) | 100 | 88.9 |
| Drawing a fork(DX) | 92.3 | 100 |
| Drawing a tick(DT) | 100 | 100 |
| Drawing a circle(DC) | 93.8 | 93.8 |
| Hand Clap(HCp) | 100 | 100 |
| Two Hand Wave(HW) | 100 | 100 |
| Punch from the side(SB) | 86.7 | 100 |
| Bending down(B) | 58.3 | 77.8 |
| Kick Forward(FK) | 100 | 100 |
| Kick Side(SK) | 100 | 100 |
| Jogging(J) | 100 | 100 |
| Swing tennis racket(TSw) | 93.8 | 83.3 |
| Overhand serve(TSr) | 100 | 93.8 |
| Swing a golf club(GS) | 100 | 100 |
| Picking up and throwing(PT) | 75 | 75 |
| Overall recognition accuracy rate | 94.88 | 93.31 |
| The number of joint-points | accuracy rate (%) | |||
|---|---|---|---|---|
| 20 | 23 | 25 | ||
| ✓ | ✓ | 0.1 | 94.09 | |
| 0.2 | 94.28 | |||
| 0.3 | 93.70 | |||
| ✓ | ✓ | 0.1 | 93.31 | |
| 0.2 | 94.88 | |||
| 0.3 | 92.92 | |||
| ✓ | ✓ | ✓ | 0.1 | 95.67 |
| 0.2 | 92.92 | |||
| 0.3 | 93.70 | |||
| method | accuracy rate (%) |
|---|---|
| Yang et al. [20] | 93.63 |
| adaptive skeleton center point [21] | 88.47 |
| Agahian et al. [22] | 91.90 |
| Zhao et al. [10] | 94.50 |
| STACOG [23] | 93.40 |
| Zhang et al. [24] | 94.81 |
| Wu et al. [25] | 95.18 |
| You et al. [26] | 91.91 |
| Ours | 95.67 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).