Swin-Fusion: Swin-Transformer with Feature Fusion for Human Action Recognition.
Saved in:
| Title: | Swin-Fusion: Swin-Transformer with Feature Fusion for Human Action Recognition. |
|---|---|
| Authors: | Chen, Tiansheng1 (AUTHOR), Mo, Lingfei1 (AUTHOR) lfmo@seu.edu.cn |
| Source: | Neural Processing Letters. Dec2023, Vol. 55 Issue 8, p11109-11130. 22p. |
| Subjects: | Convolutional neural networks, Feature extraction, Transformer models, Image recognition (Computer vision), Human activity recognition, Computer vision |
| Abstract: | Human action recognition based on still images is one of the most challenging computer vision tasks. In the past decade, convolutional neural networks (CNNs) have developed rapidly and achieved good performance in human action recognition tasks based on still images. Due to the absence of the remote perception ability of CNNs, it is challenging to have a global structural understanding of human behavior and the overall relationship between the behavior and the environment. Recently, transformer-based models have been making a splash in computer vision, even reaching SOTA in several vision tasks. We explore the transformer's capability in human action recognition based on still images and add a simple but effective feature fusion module based on the Swin-Transformer model. More specifically, we propose a new transformer-based model for behavioral feature extraction that uses a pre-trained Swin-Transformer as the backbone network. Swin-Transformer's distinctive hierarchical structure, combined with the feature fusion module, is used to extract and fuse multi-scale behavioral information. Extensive experiments were conducted on five still image-based human action recognition datasets, including the Li's action dataset, the Stanford-40 dataset, the PPMI-24 dataset, the AUC-V1 dataset, and the AUC-V2 dataset. Results indicate that our proposed Swin-Fusion model achieves better behavior recognition than previously improved CNN-based models by sharing and reusing feature maps of different scales at multiple stages, without modifying the original backbone training method and with only increasing training resources by 1.6%. The code and models will be available at https://github.com/cts4444/Swin-Fusion. [ABSTRACT FROM AUTHOR] |
| Copyright of Neural Processing Letters is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 173763249 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Swin-Fusion: Swin-Transformer with Feature Fusion for Human Action Recognition. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Chen%2C+Tiansheng%22">Chen, Tiansheng</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Mo%2C+Lingfei%22">Mo, Lingfei</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> lfmo@seu.edu.cn</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Neural+Processing+Letters%22">Neural Processing Letters</searchLink>. Dec2023, Vol. 55 Issue 8, p11109-11130. 22p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Convolutional+neural+networks%22">Convolutional neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Feature+extraction%22">Feature extraction</searchLink><br /><searchLink fieldCode="DE" term="%22Transformer+models%22">Transformer models</searchLink><br /><searchLink fieldCode="DE" term="%22Image+recognition+%28Computer+vision%29%22">Image recognition (Computer vision)</searchLink><br /><searchLink fieldCode="DE" term="%22Human+activity+recognition%22">Human activity recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+vision%22">Computer vision</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Human action recognition based on still images is one of the most challenging computer vision tasks. In the past decade, convolutional neural networks (CNNs) have developed rapidly and achieved good performance in human action recognition tasks based on still images. Due to the absence of the remote perception ability of CNNs, it is challenging to have a global structural understanding of human behavior and the overall relationship between the behavior and the environment. Recently, transformer-based models have been making a splash in computer vision, even reaching SOTA in several vision tasks. We explore the transformer's capability in human action recognition based on still images and add a simple but effective feature fusion module based on the Swin-Transformer model. More specifically, we propose a new transformer-based model for behavioral feature extraction that uses a pre-trained Swin-Transformer as the backbone network. Swin-Transformer's distinctive hierarchical structure, combined with the feature fusion module, is used to extract and fuse multi-scale behavioral information. Extensive experiments were conducted on five still image-based human action recognition datasets, including the Li's action dataset, the Stanford-40 dataset, the PPMI-24 dataset, the AUC-V1 dataset, and the AUC-V2 dataset. Results indicate that our proposed Swin-Fusion model achieves better behavior recognition than previously improved CNN-based models by sharing and reusing feature maps of different scales at multiple stages, without modifying the original backbone training method and with only increasing training resources by 1.6%. The code and models will be available at https://github.com/cts4444/Swin-Fusion. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Neural Processing Letters is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=173763249 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s11063-023-11367-1 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 22 StartPage: 11109 Subjects: – SubjectFull: Convolutional neural networks Type: general – SubjectFull: Feature extraction Type: general – SubjectFull: Transformer models Type: general – SubjectFull: Image recognition (Computer vision) Type: general – SubjectFull: Human activity recognition Type: general – SubjectFull: Computer vision Type: general Titles: – TitleFull: Swin-Fusion: Swin-Transformer with Feature Fusion for Human Action Recognition. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Chen, Tiansheng – PersonEntity: Name: NameFull: Mo, Lingfei IsPartOfRelationships: – BibEntity: Dates: – D: 20 M: 12 Text: Dec2023 Type: published Y: 2023 Identifiers: – Type: issn-print Value: 13704621 Numbering: – Type: volume Value: 55 – Type: issue Value: 8 Titles: – TitleFull: Neural Processing Letters Type: main |
| ResultId | 1 |