Pooling in image representation: The visual codeword point of view

Saved in:
Bibliographic Details
Title: Pooling in image representation: The visual codeword point of view
Authors: Avila, Sandra1,2 sandra@dcc.ufmg.br, Thome, Nicolas1, Cord, Matthieu1, Valle, Eduardo3, de A. Araújo, Arnaldo2
Source: Computer Vision & Image Understanding. May2013, Vol. 117 Issue 5, p453-465. 13p.
Subjects: Digital image processing, Time code (Audiovisual technology), Image converters, Mathematical models, Information theory, Feature extraction
Abstract: Abstract: In this work, we propose BossaNova, a novel representation for content-based concept detection in images and videos, which enriches the Bag-of-Words model. Relying on the quantization of highly discriminant local descriptors by a codebook, and the aggregation of those quantized descriptors into a single pooled feature vector, the Bag-of-Words model has emerged as the most promising approach for concept detection on visual documents. BossaNova enhances that representation by keeping a histogram of distances between the descriptors found in the image and those in the codebook, preserving thus important information about the distribution of the local descriptors around each codeword. Contrarily to other approaches found in the literature, the non-parametric histogram representation is compact and simple to compute. BossaNova compares well with the state-of-the-art in several standard datasets: MIRFLICKR, ImageCLEF 2011, PASCAL VOC 2007 and 15-Scenes, even without using complex combinations of different local descriptors. It also complements well the cutting-edge Fisher Vector descriptors, showing even better results when employed in combination with them. BossaNova also shows good results in the challenging real-world application of pornography detection. [Copyright &y& Elsevier]
Copyright of Computer Vision & Image Understanding is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 86025214
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Pooling in image representation: The visual codeword point of view
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Avila%2C+Sandra%22">Avila, Sandra</searchLink><relatesTo>1,2</relatesTo><i> sandra@dcc.ufmg.br</i><br /><searchLink fieldCode="AR" term="%22Thome%2C+Nicolas%22">Thome, Nicolas</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Cord%2C+Matthieu%22">Cord, Matthieu</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Valle%2C+Eduardo%22">Valle, Eduardo</searchLink><relatesTo>3</relatesTo><br /><searchLink fieldCode="AR" term="%22de+A%2E+Araújo%2C+Arnaldo%22">de A. Araújo, Arnaldo</searchLink><relatesTo>2</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Computer+Vision+%26+Image+Understanding%22">Computer Vision & Image Understanding</searchLink>. May2013, Vol. 117 Issue 5, p453-465. 13p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Digital+image+processing%22">Digital image processing</searchLink><br /><searchLink fieldCode="DE" term="%22Time+code+%28Audiovisual+technology%29%22">Time code (Audiovisual technology)</searchLink><br /><searchLink fieldCode="DE" term="%22Image+converters%22">Image converters</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematical+models%22">Mathematical models</searchLink><br /><searchLink fieldCode="DE" term="%22Information+theory%22">Information theory</searchLink><br /><searchLink fieldCode="DE" term="%22Feature+extraction%22">Feature extraction</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Abstract: In this work, we propose BossaNova, a novel representation for content-based concept detection in images and videos, which enriches the Bag-of-Words model. Relying on the quantization of highly discriminant local descriptors by a codebook, and the aggregation of those quantized descriptors into a single pooled feature vector, the Bag-of-Words model has emerged as the most promising approach for concept detection on visual documents. BossaNova enhances that representation by keeping a histogram of distances between the descriptors found in the image and those in the codebook, preserving thus important information about the distribution of the local descriptors around each codeword. Contrarily to other approaches found in the literature, the non-parametric histogram representation is compact and simple to compute. BossaNova compares well with the state-of-the-art in several standard datasets: MIRFLICKR, ImageCLEF 2011, PASCAL VOC 2007 and 15-Scenes, even without using complex combinations of different local descriptors. It also complements well the cutting-edge Fisher Vector descriptors, showing even better results when employed in combination with them. BossaNova also shows good results in the challenging real-world application of pornography detection. [Copyright &y& Elsevier]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Computer Vision & Image Understanding is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=86025214
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.cviu.2012.09.007
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 13
        StartPage: 453
    Subjects:
      – SubjectFull: Digital image processing
        Type: general
      – SubjectFull: Time code (Audiovisual technology)
        Type: general
      – SubjectFull: Image converters
        Type: general
      – SubjectFull: Mathematical models
        Type: general
      – SubjectFull: Information theory
        Type: general
      – SubjectFull: Feature extraction
        Type: general
    Titles:
      – TitleFull: Pooling in image representation: The visual codeword point of view
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Avila, Sandra
      – PersonEntity:
          Name:
            NameFull: Thome, Nicolas
      – PersonEntity:
          Name:
            NameFull: Cord, Matthieu
      – PersonEntity:
          Name:
            NameFull: Valle, Eduardo
      – PersonEntity:
          Name:
            NameFull: de A. Araújo, Arnaldo
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 05
              Text: May2013
              Type: published
              Y: 2013
          Identifiers:
            – Type: issn-print
              Value: 10773142
          Numbering:
            – Type: volume
              Value: 117
            – Type: issue
              Value: 5
          Titles:
            – TitleFull: Computer Vision & Image Understanding
              Type: main
ResultId 1