Collaborative multimodal feature learning for RGB-D action recognition.

Saved in:
Bibliographic Details
Title: Collaborative multimodal feature learning for RGB-D action recognition.
Authors: Kong, Jun1 (AUTHOR), Liu, Tianshan1 (AUTHOR), Jiang, Min1 (AUTHOR) minjiang@jiangnan.edu.cn
Source: Journal of Visual Communication & Image Representation. Feb2019, Vol. 59, p537-549. 13p.
Subjects: Human activity recognition, Machine learning, Human behavior
Abstract: Highlights • Our CMFL model jointly learns shared-specific features and action classifiers. • The proposed RSTPF features extract dynamic local patterns around each human joint. • The CMFL model enables the features to be optimized for classification. • The CMFL performs well even if one or two modalities are missing in the testing stage. • A max-margin framework is introduced to fuse skeleton, depth and RGB data. Abstract The emergence of cost-effective depth sensors opens up a new dimension for RGB-D based human action recognition. In this paper, we propose a collaborative multimodal feature learning (CMFL) model for human action recognition from RGB-D sequences. Specifically, we propose a robust spatio-temporal pyramid feature (RSTPF) to capture dynamic local patterns around each human joint. The proposed CMFL model fuses multimodal data (skeleton, depth and RGB), and learns action classifiers using the fused features. The original low-level feature matrices are factorized to learn shared features and modality-specific features under a supervised fashion. The shared features describe the common structures among the three modalities while the modality-specific features capture intrinsic information of each modality. We formulate shared-specific features mining and action classifiers learning in a unified max-margin framework, and solve the formulation using an iterative optimization algorithm. Experimental results on four action datasets demonstrate the efficacy of the proposed method. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Visual Communication & Image Representation is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 135379456
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Collaborative multimodal feature learning for RGB-D action recognition.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Kong%2C+Jun%22">Kong, Jun</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Liu%2C+Tianshan%22">Liu, Tianshan</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Jiang%2C+Min%22">Jiang, Min</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> minjiang@jiangnan.edu.cn</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Journal+of+Visual+Communication+%26+Image+Representation%22">Journal of Visual Communication & Image Representation</searchLink>. Feb2019, Vol. 59, p537-549. 13p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Human+activity+recognition%22">Human activity recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Human+behavior%22">Human behavior</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Highlights • Our CMFL model jointly learns shared-specific features and action classifiers. • The proposed RSTPF features extract dynamic local patterns around each human joint. • The CMFL model enables the features to be optimized for classification. • The CMFL performs well even if one or two modalities are missing in the testing stage. • A max-margin framework is introduced to fuse skeleton, depth and RGB data. Abstract The emergence of cost-effective depth sensors opens up a new dimension for RGB-D based human action recognition. In this paper, we propose a collaborative multimodal feature learning (CMFL) model for human action recognition from RGB-D sequences. Specifically, we propose a robust spatio-temporal pyramid feature (RSTPF) to capture dynamic local patterns around each human joint. The proposed CMFL model fuses multimodal data (skeleton, depth and RGB), and learns action classifiers using the fused features. The original low-level feature matrices are factorized to learn shared features and modality-specific features under a supervised fashion. The shared features describe the common structures among the three modalities while the modality-specific features capture intrinsic information of each modality. We formulate shared-specific features mining and action classifiers learning in a unified max-margin framework, and solve the formulation using an iterative optimization algorithm. Experimental results on four action datasets demonstrate the efficacy of the proposed method. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Journal of Visual Communication & Image Representation is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=135379456
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.jvcir.2019.02.013
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 13
        StartPage: 537
    Subjects:
      – SubjectFull: Human activity recognition
        Type: general
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Human behavior
        Type: general
    Titles:
      – TitleFull: Collaborative multimodal feature learning for RGB-D action recognition.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Kong, Jun
      – PersonEntity:
          Name:
            NameFull: Liu, Tianshan
      – PersonEntity:
          Name:
            NameFull: Jiang, Min
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 02
              Text: Feb2019
              Type: published
              Y: 2019
          Identifiers:
            – Type: issn-print
              Value: 10473203
          Numbering:
            – Type: volume
              Value: 59
          Titles:
            – TitleFull: Journal of Visual Communication & Image Representation
              Type: main
ResultId 1