Document Type : Original/Review Paper

Authors

Electrical and Computer Engineering Department, Semnan University, Semnan, Iran.

10.22044/jadm.2026.16786.2810

Abstract

Monitoring the daily activities of elderly individuals plays a crucial role in accident prevention, health assessment, and improving quality of life. In this paper, we propose a lightweight and efficient convolutional neural network architecture for human activity recognition based on skeletal data. Unlike conventional approaches that rely solely on absolute joint coordinates, the proposed method incorporates short- and long-term frame differences as well as spatial variations across joints to construct complementary views, thereby providing a richer spatiotemporal representation. The architecture consists of multiple convolutional blocks with residual connections, followed by global average pooling and a fully connected layer for final classification. Experimental evaluations conducted on two benchmark datasets, NTU RGB+D and ETRI-Activity3D, demonstrate that while the proposed model may achieve slightly lower accuracy compared to some state-of-the-art methods, it offers high inference speed and low computational complexity. These characteristics make it particularly suitable for real-time applications and deployment on resource-constrained devices, especially in elderly home-care environments.

Keywords

Main Subjects

[1] J. Jang, D. Kim, C. Park, M. Jang, J. Lee, and J. Kim, "ETRI-activity3D: A large-scale RGB-D dataset for robots to recognize daily activities of the elderly," in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2020: IEEE, pp. 10990–10997.
 
[2] F. Mehmood, X. Guo, E. Chen, M. A. Akbar, A. A. Khan, and S. Ullah, "Extended multi-stream temporal-attention module for skeleton-based human action recognition (HAR)," Computers in Human Behavior, vol. 163, p. 108482, 2025.
 
[3] A. M. Ahmadi, K. Kiani, R. Rastgoo, "A Transformer-Based Model for Abnormal Activity Recognition," Journal of Modeling in Engineering, vol. 22, no. 76, pp. 213-221, 2024.
 
[4] R. Rastgoo, K. Kiani, S. Escalera, "A deep generative Skeleton-based dynamic hand gesture production model," Multimedia Tools and Applications, vol. 84, pp. 48589–48608, 2025.
 
[5] R. Mohammadizand, R. Rastgoo, "Skeleton-based Sign Language Generation Using a Transformer-based Generative Model," Journal of AI and Data Mining, 2026.
 
[6] R. Rastgoo, K. Kiani, "Face recognition using fine-tuning of Deep Convolutional Neural Network and transfer learning," Journal of Modeling in Engineering, vol. 17, no. 58, pp. 103-111, 2019.
 
[7] F. Bagherzadeh, R. Rastgoo, "Deepfake image detection using a deep hybrid convolutional neural network," Journal of Modeling in Engineering, vol. 21, no. 75, pp. 19-28, 2023.
 
[8] S. Yan, Y. Xiong, and D. Lin, "Spatial temporal graph convolutional networks for skeleton-based action recognition," in Proceedings of the AAAI conference on artificial intelligence, 2018, vol. 32, no. 1.
[9] N. Raisi, M. Rezaei, and B. Masoumi, "Attention-HAR: Advanced Human Activity Recognition Using a Deep‎ Learning Model with an Integrated Attention Mechanism," Journal of AI and Data Mining, 2025.
 
[10] Z. Li, F. Li, and G. Hua, "Dynamic Graph Attention Network for Skeleton-Based Action Recognition," Applied Sciences, vol. 15, no. 9, p. 4929, 2025.
 
[11] K. Gedamu, Y. Ji, L. Gao, Y. Yang, and H. T. Shen, "Relation-mining self-attention network for skeleton-based human action recognition," Pattern Recognition, vol. 139, p. 109455, 2023.
 
[12] A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, "Ntu rgb+ d: A large scale dataset for 3d human activity analysis," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1010–1019.
 
[13] H.-g. Chi, M. H. Ha, S. Chi, S. W. Lee, Q. Huang, and K. Ramani, "Infogcn: Representation learning for human skeleton-based action recognition," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 20186–20196.
[14] Z. Liu, H. Zhang, Z. Chen, Z. Wang, and W. Ouyang, "Disentangling and unifying graph convolutions for skeleton-based action recognition," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 143–152.
 
[15] Z. Chen, S. Li, B. Yang, Q. Li, and H. Liu, "Multi-scale spatial temporal graph convolutional network for skeleton-based action recognition," in Proceedings of the AAAI conference on artificial intelligence, 2021, vol. 35, no. 2, pp. 1113–1122.
 
[16] T. Chen et al., "Learning multi-granular spatio-temporal graph network for skeleton-based action recognition," in Proceedings of the 29th ACM international conference on multimedia, 2021, pp. 4334–4342.
 
[17] H. Liu, Y. Liu, Y. Chen, C. Yuan, B. Li, and W. Hu, "TranSkeleton: Hierarchical spatial–temporal transformer for skeleton-based action recognition," IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 8, pp. 4137–4148, 2023.
 
[18] C. Plizzari, M. Cannici, and M. Matteucci, "Spatial temporal transformer network for skeleton-based action recognition," in international conference on pattern recognition, 2021: Springer, pp. 694–701.
 
[19] H. Qiu, B. Hou, B. Ren, and X. Zhang, "Spatio-temporal tuples transformer for skeleton-based action recognition," arXiv preprint arXiv:2201.02849, 2022.
 
[20] K. Kiani, S. Rezaeirad, "A new ergodic HMM-based face recognition using DWT and half of the face," in 5th Conference on Knowledge Based Engineering and Innovation (KBEI), 2019, pp. 531-536.
 
[21] J. Do and M. Kim, "Skateformer: skeletal-temporal transformer for human action recognition," in European Conference on Computer Vision, 2024: Springer, pp. 401–420.
[22] J. Zhang, Y. Jia, W. Xie, and Z. Tu, "Zoom transformer for skeleton-based group activity recognition," IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 12, pp. 8646–8659, 2022.
 
[23] W. Wu, C. Zheng, Z. Yang, C. Chen, S. Das, and A. Lu, "Frequency guidance matters: Skeletal action recognition by frequency-aware mixed transformer," in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 4660–4669.
 
[24] Y. Zhou et al., "Hypergraph transformer for skeleton-based action recognition," arXiv preprint arXiv:2211.09590, 2022.
 
[25] K. Cheng, Y. Zhang, X. He, W. Chen, J. Cheng, and H. Lu, "Skeleton-based action recognition with shift graph convolutional network," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 183–192.
 
[26] X. Shu, J. Yang, R. Yan, and Y. Song, "Expansion-squeeze-excitation fusion network for elderly activity recognition," IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 8, pp. 5281–5292, 2022.