Multimodal Adaptive Graph Convolution: Toward Robust Skeleton-based Action Recognition.

Saved in:
Bibliographic Details
Title: Multimodal Adaptive Graph Convolution: Toward Robust Skeleton-based Action Recognition.
Authors: Guo, Lijuan1 guolj.sy@gx.csg.cn, Xie, Guoshan2 xiegs.sy@gx.csg.cn, Mo, Jing3 moj@gx.csg.cn, Shi, Fengwei4 shifw.lbg@gx.csg.cn, Wang, Le2 wangl.sy@gx.csg.cn
Source: IAENG International Journal of Computer Science. Jun2026, Vol. 53 Issue 6, p2400-2409. 10p.
Subjects: Human activity recognition, Graph neural networks, Learning, Multisensor data fusion, Spatiotemporal processes
Abstract: Significant progress has been made in human action recognition through the use of various modalities, such as RGB videos, depth maps, skeleton data, and infrared imaging. Skeleton-based methods that use graph convolutional networks (GCNs) are particularly impressive. Recent methods employ deformable graph convolution operations to automatically identify semantically important body joints, increasing computational efficiency while maintaining high recognition accuracy. However, current methods still have three basic limitations: (1) rigid feature extraction that cannot adaptively learn discriminative joint-level representations, (2) overreliance on fixed hyperparameters, and (3) inadequate ability to model the complex spatiotemporal variations in skeleton sequences. These deficiencies significantly limit their generalizability across diverse datasets. To overcome these challenges, we propose an innovative multimodal adaptive graph convolution network (MMA-GCN), which can effectively improve skeleton-based action recognition. First, we develop an adaptive parameterization mechanism that dynamically adjusts network parameters to increase generalization robustness. Second, our model automatically learns deformable temporal sampling patterns while adaptively optimizing spatial correlation weights, enabling context-sensitive perception of discriminative features. Third, we propose a principled multimodal fusion strategy that effectively integrates four complementary representations: joint positions, bone vectors, motion velocities, and joint-bone correlations. This comprehensive representation captures richer action semantics than conventional single-modal approaches do. Through extensive experiments on two large-scale benchmarks (NTURGBD60 and NTURGBD120), we demonstrate that our method achieves new state-of-the-art performance. More importantly, compared with existing approaches, the proposed techniques substantially improve cross-dataset generalization, validating the effectiveness of our architectural innovations. The excellent results stem from the unique ability of our method to adaptively learn both spatial and temporal features while exploiting multimodal skeletal information. [ABSTRACT FROM AUTHOR]
Copyright of IAENG International Journal of Computer Science is the property of International Association of Engineers (IAENG) and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:Significant progress has been made in human action recognition through the use of various modalities, such as RGB videos, depth maps, skeleton data, and infrared imaging. Skeleton-based methods that use graph convolutional networks (GCNs) are particularly impressive. Recent methods employ deformable graph convolution operations to automatically identify semantically important body joints, increasing computational efficiency while maintaining high recognition accuracy. However, current methods still have three basic limitations: (1) rigid feature extraction that cannot adaptively learn discriminative joint-level representations, (2) overreliance on fixed hyperparameters, and (3) inadequate ability to model the complex spatiotemporal variations in skeleton sequences. These deficiencies significantly limit their generalizability across diverse datasets. To overcome these challenges, we propose an innovative multimodal adaptive graph convolution network (MMA-GCN), which can effectively improve skeleton-based action recognition. First, we develop an adaptive parameterization mechanism that dynamically adjusts network parameters to increase generalization robustness. Second, our model automatically learns deformable temporal sampling patterns while adaptively optimizing spatial correlation weights, enabling context-sensitive perception of discriminative features. Third, we propose a principled multimodal fusion strategy that effectively integrates four complementary representations: joint positions, bone vectors, motion velocities, and joint-bone correlations. This comprehensive representation captures richer action semantics than conventional single-modal approaches do. Through extensive experiments on two large-scale benchmarks (NTURGBD60 and NTURGBD120), we demonstrate that our method achieves new state-of-the-art performance. More importantly, compared with existing approaches, the proposed techniques substantially improve cross-dataset generalization, validating the effectiveness of our architectural innovations. The excellent results stem from the unique ability of our method to adaptively learn both spatial and temporal features while exploiting multimodal skeletal information. [ABSTRACT FROM AUTHOR]
ISSN:1819656X