2609005192
  • Open Access
  • Article

MG2CL: Multi-Granularity Graph Contrastive Learning for Skeleton-Based Action Recognition

  • Zhize Wu 1,2,   
  • Pensong Wang 1,   
  • Xin Chen 1,   
  • Thomas Weise 1,   
  • Fei Liu 3,4,   
  • Shengwei Ji 1,4,*

Received: 04 Jun 2026 | Revised: 05 Sep 2026 | Accepted: 16 Sep 2026 | Published: 22 Sep 2026

Abstract

Skeleton-based action recognition has gained increasing attention as a robust alternative to RGB-based methods due to its resilience to background clutter and lighting variations. Existing approaches have achieved notable progress, particularly with graph-based models that capture topological and temporal patterns of human joints. However, most methods still struggle to model relationships between non-naturally connected joints and fail to fully exploit the fine-grained spatio-temporal characteristics of skeleton sequences, limiting generalization and recognition accuracy. To address these limitations, we propose MG2CL, a Multi-Granularity Graph Contrastive Learning framework designed to enhance skeleton representation learning. Our approach employs a multi-granularity graph structure to model both short- and long-range joint interactions through complementary body-part relations. Furthermore, we introduce a novel spatio-temporal masking strategy for data augmentation, encouraging the model to learn more diverse and informative patterns. A semantic-level memory bank, built upon the multi-granularity graph, is integrated to reinforce the model’s ability to distinguish subtle action variations. Extensive experiments show that MG2CL achieves competitive performance on widely used benchmark datasets, including NTU RGB+D, NTU RGB+D 120, and Northwestern-UCLA. These results demonstrate the framework’s strong generalization and discriminative capability. Our findings suggest that incorporating multi-granularity structural information and contrastive learning principles can lead to more robust and flexible skeleton-based action models. MG2CL offers a promising direction for future work in representation learning for spatio-temporal graph data, with potential applications in human-computer interaction, surveillance, and healthcare.

References 

  • 1.

    Khaleghi, L.; Sepas-Moghaddam, A.; Marshall, J.; et al. Multiview Video-Based 3-D Hand Pose Estimation. IEEE Trans. Artif. Intell. 2023, 4, 896–909.

  • 2.

    Cao, C.; Wang, Y.; Zhang, Y.; et al. Co-Occurrence Matters: Learning Action Relation for Temporal Action Localization. IEEE Trans. Circuits Syst. Video Technol. 2024, 34, 3327–3339.

  • 3.

    Vahdani, E.; Tian, Y. Deep Learning-Based Action Detection in Untrimmed Videos: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 4302–4320.

  • 4.

    Sun, Z.; Ke, Q.; Rahmani, H.; et al. Human Action Recognition From Various Data Modalities: A Review. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 3200–3225.

  • 5.

    Ahmad, T.; Jin, L.; Zhang, X.; et al. Graph Convolutional Neural Network for Human Action Recognition: A Comprehensive Survey. IEEE Trans. Artif. Intell. 2021, 2, 128–145.

  • 6.

    Yan, H.; Liu, Y.; Wei, Y.; et al. SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-training. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023.

  • 7.

    Do, J.; Kim, M. SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024.

  • 8.

    Liu, J.; Wang, X.; Wang, C.; et al. Temporal Decoupling Graph Convolutional Network for Skeleton-Based Gesture Recognition. IEEE Trans. Multimed. 2024, 26, 811–823.

  • 9.

    Yan, S.; Xiong, Y.; Lin, D. Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence (AAAI’18), New Orleans, LA, USA, 2–7 February 2018; pp. 7444–7452.

  • 10.

    Zhou, M.; Lu, J.; Zhang, G. Multi-Scale Adaptive Convolutional Graph for Multi-Stream Concept Drift. IEEE Trans. Knowl. Data Eng. 2026, 38, 2982–2994.

  • 11.

    Zhu, Y.; Shuai, H.; Liu, G.; et al. Multilevel Spatial–Temporal Excited Graph Network for Skeleton-Based Action Recognition. IEEE Trans. Image Process. 2023, 32, 496–508.

  • 12.

    Lee, J.; Lee, M.; Lee, D.; et al. Hierarchically decomposed graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 10444–10453.

  • 13.

    Huang, H.; Zhou, H.; Wang, J. Graph Contrastive Learning for Skeleton-based Action Recognition. In Proceedings of the International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023.

  • 14.

    Pang, C.; Lu, X.; Lyu, L. Skeleton-based action recognition through contrasting two-stream spatial-temporal networks. IEEE Trans. Multimed. 2023, 25, 8699–8711.

  • 15.

    Wu, Z.; Sun, P.; Chen, X.; et al. SelfGCN: Graph Convolution Network with Self-Attention for Skeleton-Based Action Recognition. IEEE Trans. Image Process. 2024, 33, 4391–4403.

  • 16.

    Tian, H.; Ma, X.; Li, X.; et al. Skeleton-Based Action Recognition with Select-Assemble-Normalize Graph Convolutional Networks. IEEE Trans. Multimed. 2023, 25, 8527–8538.

  • 17.

    Cormier, M.; Schmid, Y.; Beyerer, J. Enhancing Skeleton-Based Action Recognition in Real-World Scenarios through Realistic Data Augmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 3–8 January 2024; pp. 290–299.

  • 18.

    Ni, B.; Paramathayalan, V.R.; Moulin, P. Multiple granularity analysis for fine-grained action detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 756–763.

  • 19.

    Chen, T.; Zhou, D.; Wang, J.; et al. Learning multi-granular spatio-temporal graph network for skeleton-based action recognition. In Proceedings of the 29th ACM International Conference on Multimedia, Virtual, 20–24 October 2021; pp. 4334–4342.

  • 20.

    Chen, H.; Shen, Y.; Zhang, Y.; et al. Skeleton-based action recognition through dual-granularity feature fusion with self-adapting graph convolution and multi-scale temporal convolution. Neurocomputing 2025, 639, 130261.

  • 21.

    Shahroudy, A.; Liu, J.; Ng, T.; et al. NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR’16), Las Vegas, NV, USA, 26 June–1 July 2016; pp. 1010–1019.

  • 22.

    Liu, J.; Shahroudy, A.; Perez, M.; et al. NTU RGB+D 120: A large-scale benchmark for 3D human activity understanding. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 42, 2684–2701.

  • 23.

    Wang, J.; Nie, X.; Xia, Y.; et al. Cross-view action modeling, learning and recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 2649–2656.

  • 24.

    Liu, Z.; Zhang, H.; Chen, Z.; et al. Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR’20), Virtual, 14–19 June 2020; pp. 140–149.

  • 25.

    Chen, Y.; Zhang, Z.; Yuan, C.; et al. Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV’21), Virtual, 11–17 October 2021; pp. 13339–13348.

  • 26.

    Wu, Z.; Ding, Y.; Wan, L.; et al. Local and global self-attention enhanced graph convolutional network for skeleton-based action recognition. Pattern Recognit. 2025, 159, 111106.

  • 27.

    Myung, W.; Su, N.; Xue, J.H.; et al. DeGCN: Deformable Graph Convolutional Networks for Skeleton-Based Action Recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 288–300.

  • 28.

    Huang, L.; Huang, Y.; Ouyang, W.; et al. Part-level graph convolutional network for skeleton-based action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA, 7–12 February 2020; Volume 34, pp. 11045–11052.

  • 29.

    Trivedi, N.; Sarvadevabhatla, K. PSUMNet: Unified modality part streams are all you need for efficient pose-based action recognition. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 211–227.

  • 30.

    Li, L.; Wang, M.; Ni, B. 3D human action representation learning via cross-view consistency pursuit. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 4741–4750.

  • 31.

    Wu, C.; Wu, J. SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action Recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February 2024; pp. 5949–5957.

  • 32.

    Qin, Z.; Liu, Y.; Ji, P.; et al. Fusing Higher-Order Features in Graph Neural Networks for Skeleton-Based Action Recognition. IEEE Trans. Neural Netw. Learn. Syst. 2024, 35, 4783–4797.

  • 33.

    Xiang, L.; Wang, Z. Joint mixing data augmentation for skeleton-based action recognition. ACM Trans. Multimed. Comput. Commun. Appl. 2025, 21, 108.

  • 34.

    Song, Y.F.; Zhang, Z.; Shan, C.; et al. Constructing stronger and faster baselines for skeleton-based action recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 1474–1488.

  • 35.

    Shi, L.; Zhang, Y.; Cheng, J.; et al. Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR’19), Long Beach, CA, USA, 15–20 June 2019; pp. 12026–12035.

  • 36.

    Chen, Z.; Li, S.; Yang, B.; et al. Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action Recognition. In Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence (AAAI’21), Virtual, 2–9 February 2021; pp. 1113–1122.

  • 37.

    Chi, H.G.; Ha, M.H.; Chi, S.; et al. InfoGCN: Representation learning for human skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 20186–20196.

  • 38.

    Zhuang, T.; Qin, Z.; Ding, Y.; et al. Temporal Refinement Graph Convolutional Network for Skeleton-Based Action Recognition. IEEE Trans. Artif. Intell. 2024, 5, 1586–1598.

Share this article:
How to Cite
Wu, Z.; Wang, P.; Chen, X.; Weise, T.; Liu, F.; Ji, S. MG2CL: Multi-Granularity Graph Contrastive Learning for Skeleton-Based Action Recognition. Transactions on Graph Intelligence and Network Applications 2026. https://doi.org/10.53941/tgina.2026.100007.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.
Article Metrics
34
Article Views
0
Citations