Collaborative multimodal feature learning for RGB-D action recognition.
Saved in:
| Title: | Collaborative multimodal feature learning for RGB-D action recognition. |
|---|---|
| Authors: | Kong, Jun1 (AUTHOR), Liu, Tianshan1 (AUTHOR), Jiang, Min1 (AUTHOR) minjiang@jiangnan.edu.cn |
| Source: | Journal of Visual Communication & Image Representation. Feb2019, Vol. 59, p537-549. 13p. |
| Subjects: | Human activity recognition, Machine learning, Human behavior |
| Abstract: | Highlights • Our CMFL model jointly learns shared-specific features and action classifiers. • The proposed RSTPF features extract dynamic local patterns around each human joint. • The CMFL model enables the features to be optimized for classification. • The CMFL performs well even if one or two modalities are missing in the testing stage. • A max-margin framework is introduced to fuse skeleton, depth and RGB data. Abstract The emergence of cost-effective depth sensors opens up a new dimension for RGB-D based human action recognition. In this paper, we propose a collaborative multimodal feature learning (CMFL) model for human action recognition from RGB-D sequences. Specifically, we propose a robust spatio-temporal pyramid feature (RSTPF) to capture dynamic local patterns around each human joint. The proposed CMFL model fuses multimodal data (skeleton, depth and RGB), and learns action classifiers using the fused features. The original low-level feature matrices are factorized to learn shared features and modality-specific features under a supervised fashion. The shared features describe the common structures among the three modalities while the modality-specific features capture intrinsic information of each modality. We formulate shared-specific features mining and action classifiers learning in a unified max-margin framework, and solve the formulation using an iterative optimization algorithm. Experimental results on four action datasets demonstrate the efficacy of the proposed method. [ABSTRACT FROM AUTHOR] |
| Copyright of Journal of Visual Communication & Image Representation is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 135379456 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Collaborative multimodal feature learning for RGB-D action recognition. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Kong%2C+Jun%22">Kong, Jun</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Liu%2C+Tianshan%22">Liu, Tianshan</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Jiang%2C+Min%22">Jiang, Min</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> minjiang@jiangnan.edu.cn</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+Visual+Communication+%26+Image+Representation%22">Journal of Visual Communication & Image Representation</searchLink>. Feb2019, Vol. 59, p537-549. 13p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Human+activity+recognition%22">Human activity recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Human+behavior%22">Human behavior</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Highlights • Our CMFL model jointly learns shared-specific features and action classifiers. • The proposed RSTPF features extract dynamic local patterns around each human joint. • The CMFL model enables the features to be optimized for classification. • The CMFL performs well even if one or two modalities are missing in the testing stage. • A max-margin framework is introduced to fuse skeleton, depth and RGB data. Abstract The emergence of cost-effective depth sensors opens up a new dimension for RGB-D based human action recognition. In this paper, we propose a collaborative multimodal feature learning (CMFL) model for human action recognition from RGB-D sequences. Specifically, we propose a robust spatio-temporal pyramid feature (RSTPF) to capture dynamic local patterns around each human joint. The proposed CMFL model fuses multimodal data (skeleton, depth and RGB), and learns action classifiers using the fused features. The original low-level feature matrices are factorized to learn shared features and modality-specific features under a supervised fashion. The shared features describe the common structures among the three modalities while the modality-specific features capture intrinsic information of each modality. We formulate shared-specific features mining and action classifiers learning in a unified max-margin framework, and solve the formulation using an iterative optimization algorithm. Experimental results on four action datasets demonstrate the efficacy of the proposed method. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Journal of Visual Communication & Image Representation is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=135379456 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1016/j.jvcir.2019.02.013 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 13 StartPage: 537 Subjects: – SubjectFull: Human activity recognition Type: general – SubjectFull: Machine learning Type: general – SubjectFull: Human behavior Type: general Titles: – TitleFull: Collaborative multimodal feature learning for RGB-D action recognition. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Kong, Jun – PersonEntity: Name: NameFull: Liu, Tianshan – PersonEntity: Name: NameFull: Jiang, Min IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 02 Text: Feb2019 Type: published Y: 2019 Identifiers: – Type: issn-print Value: 10473203 Numbering: – Type: volume Value: 59 Titles: – TitleFull: Journal of Visual Communication & Image Representation Type: main |
| ResultId | 1 |