ViDSOD-100: A New Dataset and a Baseline Model for RGB-D Video Salient Object Detection.

Saved in:
Bibliographic Details
Title: ViDSOD-100: A New Dataset and a Baseline Model for RGB-D Video Salient Object Detection.
Authors: Lin, Junhao1 (AUTHOR), Zhu, Lei1,2 (AUTHOR) leizhu@ust.hk, Shen, Jiaxing3 (AUTHOR), Fu, Huazhu4 (AUTHOR), Zhang, Qing5 (AUTHOR), Wang, Liansheng2 (AUTHOR)
Source: International Journal of Computer Vision. Nov2024, Vol. 132 Issue 11, p5173-5191. 19p.
Subjects: Object recognition (Computer vision), Data release, Videos, Detectors, Annotations
Abstract: With the rapid development of depth sensor, more and more RGB-D videos could be obtained. Identifying the foreground in RGB-D videos is a fundamental and important task. However, the existing salient object detection (SOD) works only focus on either static RGB-D images or RGB videos, ignoring the collaborating of RGB-D and video information. In this paper, we first collect a new annotated RGB-D video SOD (ViDSOD-100) dataset, which contains 100 videos within a total of 9362 frames, acquired from diverse natural scenes. All the frames in each video are manually annotated to a high-quality saliency annotation. Moreover, we propose a new baseline model, named attentive triple-fusion network (ATF-Net), for RGB-D video salient object detection. Our method aggregates the appearance information from an input RGB image, spatio-temporal information from an estimated motion map, and the geometry information from the depth map by devising three modality-specific branches and a multi-modality integration branch. The modality-specific branches extract the representation of different inputs, while the multi-modality integration branch combines the multi-level modality-specific features by introducing the encoder feature aggregation (MEA) modules and decoder feature aggregation (MDA) modules. The experimental findings conducted on both our newly introduced ViDSOD-100 dataset and the well-established DAVSOD dataset highlight the superior performance of the proposed ATF-Net.This performance enhancement is demonstrated both quantitatively and qualitatively, surpassing the capabilities of current state-of-the-art techniques across various domains, including RGB-D saliency detection, video saliency detection, and video object segmentation. We shall release our data, our results, and our code upon the publication of this work. [ABSTRACT FROM AUTHOR]
Copyright of International Journal of Computer Vision is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 180501479
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: ViDSOD-100: A New Dataset and a Baseline Model for RGB-D Video Salient Object Detection.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Lin%2C+Junhao%22">Lin, Junhao</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Zhu%2C+Lei%22">Zhu, Lei</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<i> leizhu@ust.hk</i><br /><searchLink fieldCode="AR" term="%22Shen%2C+Jiaxing%22">Shen, Jiaxing</searchLink><relatesTo>3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Fu%2C+Huazhu%22">Fu, Huazhu</searchLink><relatesTo>4</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Zhang%2C+Qing%22">Zhang, Qing</searchLink><relatesTo>5</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Wang%2C+Liansheng%22">Wang, Liansheng</searchLink><relatesTo>2</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22International+Journal+of+Computer+Vision%22">International Journal of Computer Vision</searchLink>. Nov2024, Vol. 132 Issue 11, p5173-5191. 19p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Object+recognition+%28Computer+vision%29%22">Object recognition (Computer vision)</searchLink><br /><searchLink fieldCode="DE" term="%22Data+release%22">Data release</searchLink><br /><searchLink fieldCode="DE" term="%22Videos%22">Videos</searchLink><br /><searchLink fieldCode="DE" term="%22Detectors%22">Detectors</searchLink><br /><searchLink fieldCode="DE" term="%22Annotations%22">Annotations</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: With the rapid development of depth sensor, more and more RGB-D videos could be obtained. Identifying the foreground in RGB-D videos is a fundamental and important task. However, the existing salient object detection (SOD) works only focus on either static RGB-D images or RGB videos, ignoring the collaborating of RGB-D and video information. In this paper, we first collect a new annotated RGB-D video SOD (ViDSOD-100) dataset, which contains 100 videos within a total of 9362 frames, acquired from diverse natural scenes. All the frames in each video are manually annotated to a high-quality saliency annotation. Moreover, we propose a new baseline model, named attentive triple-fusion network (ATF-Net), for RGB-D video salient object detection. Our method aggregates the appearance information from an input RGB image, spatio-temporal information from an estimated motion map, and the geometry information from the depth map by devising three modality-specific branches and a multi-modality integration branch. The modality-specific branches extract the representation of different inputs, while the multi-modality integration branch combines the multi-level modality-specific features by introducing the encoder feature aggregation (MEA) modules and decoder feature aggregation (MDA) modules. The experimental findings conducted on both our newly introduced ViDSOD-100 dataset and the well-established DAVSOD dataset highlight the superior performance of the proposed ATF-Net.This performance enhancement is demonstrated both quantitatively and qualitatively, surpassing the capabilities of current state-of-the-art techniques across various domains, including RGB-D saliency detection, video saliency detection, and video object segmentation. We shall release our data, our results, and our code upon the publication of this work. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of International Journal of Computer Vision is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=180501479
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s11263-024-02051-5
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 19
        StartPage: 5173
    Subjects:
      – SubjectFull: Object recognition (Computer vision)
        Type: general
      – SubjectFull: Data release
        Type: general
      – SubjectFull: Videos
        Type: general
      – SubjectFull: Detectors
        Type: general
      – SubjectFull: Annotations
        Type: general
    Titles:
      – TitleFull: ViDSOD-100: A New Dataset and a Baseline Model for RGB-D Video Salient Object Detection.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Lin, Junhao
      – PersonEntity:
          Name:
            NameFull: Zhu, Lei
      – PersonEntity:
          Name:
            NameFull: Shen, Jiaxing
      – PersonEntity:
          Name:
            NameFull: Fu, Huazhu
      – PersonEntity:
          Name:
            NameFull: Zhang, Qing
      – PersonEntity:
          Name:
            NameFull: Wang, Liansheng
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 11
              Text: Nov2024
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-print
              Value: 09205691
          Numbering:
            – Type: volume
              Value: 132
            – Type: issue
              Value: 11
          Titles:
            – TitleFull: International Journal of Computer Vision
              Type: main
ResultId 1