Multimodal Adaptive Graph Convolution: Toward Robust Skeleton-based Action Recognition.

Saved in:
Bibliographic Details
Title: Multimodal Adaptive Graph Convolution: Toward Robust Skeleton-based Action Recognition.
Authors: Guo, Lijuan1 guolj.sy@gx.csg.cn, Xie, Guoshan2 xiegs.sy@gx.csg.cn, Mo, Jing3 moj@gx.csg.cn, Shi, Fengwei4 shifw.lbg@gx.csg.cn, Wang, Le2 wangl.sy@gx.csg.cn
Source: IAENG International Journal of Computer Science. Jun2026, Vol. 53 Issue 6, p2400-2409. 10p.
Subjects: Human activity recognition, Graph neural networks, Learning, Multisensor data fusion, Spatiotemporal processes
Abstract: Significant progress has been made in human action recognition through the use of various modalities, such as RGB videos, depth maps, skeleton data, and infrared imaging. Skeleton-based methods that use graph convolutional networks (GCNs) are particularly impressive. Recent methods employ deformable graph convolution operations to automatically identify semantically important body joints, increasing computational efficiency while maintaining high recognition accuracy. However, current methods still have three basic limitations: (1) rigid feature extraction that cannot adaptively learn discriminative joint-level representations, (2) overreliance on fixed hyperparameters, and (3) inadequate ability to model the complex spatiotemporal variations in skeleton sequences. These deficiencies significantly limit their generalizability across diverse datasets. To overcome these challenges, we propose an innovative multimodal adaptive graph convolution network (MMA-GCN), which can effectively improve skeleton-based action recognition. First, we develop an adaptive parameterization mechanism that dynamically adjusts network parameters to increase generalization robustness. Second, our model automatically learns deformable temporal sampling patterns while adaptively optimizing spatial correlation weights, enabling context-sensitive perception of discriminative features. Third, we propose a principled multimodal fusion strategy that effectively integrates four complementary representations: joint positions, bone vectors, motion velocities, and joint-bone correlations. This comprehensive representation captures richer action semantics than conventional single-modal approaches do. Through extensive experiments on two large-scale benchmarks (NTURGBD60 and NTURGBD120), we demonstrate that our method achieves new state-of-the-art performance. More importantly, compared with existing approaches, the proposed techniques substantially improve cross-dataset generalization, validating the effectiveness of our architectural innovations. The excellent results stem from the unique ability of our method to adaptively learn both spatial and temporal features while exploiting multimodal skeletal information. [ABSTRACT FROM AUTHOR]
Copyright of IAENG International Journal of Computer Science is the property of International Association of Engineers (IAENG) and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 194196022
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Multimodal Adaptive Graph Convolution: Toward Robust Skeleton-based Action Recognition.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Guo%2C+Lijuan%22">Guo, Lijuan</searchLink><relatesTo>1</relatesTo><i> guolj.sy@gx.csg.cn</i><br /><searchLink fieldCode="AR" term="%22Xie%2C+Guoshan%22">Xie, Guoshan</searchLink><relatesTo>2</relatesTo><i> xiegs.sy@gx.csg.cn</i><br /><searchLink fieldCode="AR" term="%22Mo%2C+Jing%22">Mo, Jing</searchLink><relatesTo>3</relatesTo><i> moj@gx.csg.cn</i><br /><searchLink fieldCode="AR" term="%22Shi%2C+Fengwei%22">Shi, Fengwei</searchLink><relatesTo>4</relatesTo><i> shifw.lbg@gx.csg.cn</i><br /><searchLink fieldCode="AR" term="%22Wang%2C+Le%22">Wang, Le</searchLink><relatesTo>2</relatesTo><i> wangl.sy@gx.csg.cn</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22IAENG+International+Journal+of+Computer+Science%22">IAENG International Journal of Computer Science</searchLink>. Jun2026, Vol. 53 Issue 6, p2400-2409. 10p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Human+activity+recognition%22">Human activity recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Graph+neural+networks%22">Graph neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Learning%22">Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Multisensor+data+fusion%22">Multisensor data fusion</searchLink><br /><searchLink fieldCode="DE" term="%22Spatiotemporal+processes%22">Spatiotemporal processes</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Significant progress has been made in human action recognition through the use of various modalities, such as RGB videos, depth maps, skeleton data, and infrared imaging. Skeleton-based methods that use graph convolutional networks (GCNs) are particularly impressive. Recent methods employ deformable graph convolution operations to automatically identify semantically important body joints, increasing computational efficiency while maintaining high recognition accuracy. However, current methods still have three basic limitations: (1) rigid feature extraction that cannot adaptively learn discriminative joint-level representations, (2) overreliance on fixed hyperparameters, and (3) inadequate ability to model the complex spatiotemporal variations in skeleton sequences. These deficiencies significantly limit their generalizability across diverse datasets. To overcome these challenges, we propose an innovative multimodal adaptive graph convolution network (MMA-GCN), which can effectively improve skeleton-based action recognition. First, we develop an adaptive parameterization mechanism that dynamically adjusts network parameters to increase generalization robustness. Second, our model automatically learns deformable temporal sampling patterns while adaptively optimizing spatial correlation weights, enabling context-sensitive perception of discriminative features. Third, we propose a principled multimodal fusion strategy that effectively integrates four complementary representations: joint positions, bone vectors, motion velocities, and joint-bone correlations. This comprehensive representation captures richer action semantics than conventional single-modal approaches do. Through extensive experiments on two large-scale benchmarks (NTURGBD60 and NTURGBD120), we demonstrate that our method achieves new state-of-the-art performance. More importantly, compared with existing approaches, the proposed techniques substantially improve cross-dataset generalization, validating the effectiveness of our architectural innovations. The excellent results stem from the unique ability of our method to adaptively learn both spatial and temporal features while exploiting multimodal skeletal information. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of IAENG International Journal of Computer Science is the property of International Association of Engineers (IAENG) and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=194196022
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 10
        StartPage: 2400
    Subjects:
      – SubjectFull: Human activity recognition
        Type: general
      – SubjectFull: Graph neural networks
        Type: general
      – SubjectFull: Learning
        Type: general
      – SubjectFull: Multisensor data fusion
        Type: general
      – SubjectFull: Spatiotemporal processes
        Type: general
    Titles:
      – TitleFull: Multimodal Adaptive Graph Convolution: Toward Robust Skeleton-based Action Recognition.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Guo, Lijuan
      – PersonEntity:
          Name:
            NameFull: Xie, Guoshan
      – PersonEntity:
          Name:
            NameFull: Mo, Jing
      – PersonEntity:
          Name:
            NameFull: Shi, Fengwei
      – PersonEntity:
          Name:
            NameFull: Wang, Le
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: Jun2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 1819656X
          Numbering:
            – Type: volume
              Value: 53
            – Type: issue
              Value: 6
          Titles:
            – TitleFull: IAENG International Journal of Computer Science
              Type: main
ResultId 1