Pooling in image representation: The visual codeword point of view
Saved in:
| Title: | Pooling in image representation: The visual codeword point of view |
|---|---|
| Authors: | Avila, Sandra1,2 sandra@dcc.ufmg.br, Thome, Nicolas1, Cord, Matthieu1, Valle, Eduardo3, de A. Araújo, Arnaldo2 |
| Source: | Computer Vision & Image Understanding. May2013, Vol. 117 Issue 5, p453-465. 13p. |
| Subjects: | Digital image processing, Time code (Audiovisual technology), Image converters, Mathematical models, Information theory, Feature extraction |
| Abstract: | Abstract: In this work, we propose BossaNova, a novel representation for content-based concept detection in images and videos, which enriches the Bag-of-Words model. Relying on the quantization of highly discriminant local descriptors by a codebook, and the aggregation of those quantized descriptors into a single pooled feature vector, the Bag-of-Words model has emerged as the most promising approach for concept detection on visual documents. BossaNova enhances that representation by keeping a histogram of distances between the descriptors found in the image and those in the codebook, preserving thus important information about the distribution of the local descriptors around each codeword. Contrarily to other approaches found in the literature, the non-parametric histogram representation is compact and simple to compute. BossaNova compares well with the state-of-the-art in several standard datasets: MIRFLICKR, ImageCLEF 2011, PASCAL VOC 2007 and 15-Scenes, even without using complex combinations of different local descriptors. It also complements well the cutting-edge Fisher Vector descriptors, showing even better results when employed in combination with them. BossaNova also shows good results in the challenging real-world application of pornography detection. [Copyright &y& Elsevier] |
| Copyright of Computer Vision & Image Understanding is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 86025214 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Pooling in image representation: The visual codeword point of view – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Avila%2C+Sandra%22">Avila, Sandra</searchLink><relatesTo>1,2</relatesTo><i> sandra@dcc.ufmg.br</i><br /><searchLink fieldCode="AR" term="%22Thome%2C+Nicolas%22">Thome, Nicolas</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Cord%2C+Matthieu%22">Cord, Matthieu</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Valle%2C+Eduardo%22">Valle, Eduardo</searchLink><relatesTo>3</relatesTo><br /><searchLink fieldCode="AR" term="%22de+A%2E+Araújo%2C+Arnaldo%22">de A. Araújo, Arnaldo</searchLink><relatesTo>2</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Computer+Vision+%26+Image+Understanding%22">Computer Vision & Image Understanding</searchLink>. May2013, Vol. 117 Issue 5, p453-465. 13p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Digital+image+processing%22">Digital image processing</searchLink><br /><searchLink fieldCode="DE" term="%22Time+code+%28Audiovisual+technology%29%22">Time code (Audiovisual technology)</searchLink><br /><searchLink fieldCode="DE" term="%22Image+converters%22">Image converters</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematical+models%22">Mathematical models</searchLink><br /><searchLink fieldCode="DE" term="%22Information+theory%22">Information theory</searchLink><br /><searchLink fieldCode="DE" term="%22Feature+extraction%22">Feature extraction</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Abstract: In this work, we propose BossaNova, a novel representation for content-based concept detection in images and videos, which enriches the Bag-of-Words model. Relying on the quantization of highly discriminant local descriptors by a codebook, and the aggregation of those quantized descriptors into a single pooled feature vector, the Bag-of-Words model has emerged as the most promising approach for concept detection on visual documents. BossaNova enhances that representation by keeping a histogram of distances between the descriptors found in the image and those in the codebook, preserving thus important information about the distribution of the local descriptors around each codeword. Contrarily to other approaches found in the literature, the non-parametric histogram representation is compact and simple to compute. BossaNova compares well with the state-of-the-art in several standard datasets: MIRFLICKR, ImageCLEF 2011, PASCAL VOC 2007 and 15-Scenes, even without using complex combinations of different local descriptors. It also complements well the cutting-edge Fisher Vector descriptors, showing even better results when employed in combination with them. BossaNova also shows good results in the challenging real-world application of pornography detection. [Copyright &y& Elsevier] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Computer Vision & Image Understanding is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=86025214 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1016/j.cviu.2012.09.007 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 13 StartPage: 453 Subjects: – SubjectFull: Digital image processing Type: general – SubjectFull: Time code (Audiovisual technology) Type: general – SubjectFull: Image converters Type: general – SubjectFull: Mathematical models Type: general – SubjectFull: Information theory Type: general – SubjectFull: Feature extraction Type: general Titles: – TitleFull: Pooling in image representation: The visual codeword point of view Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Avila, Sandra – PersonEntity: Name: NameFull: Thome, Nicolas – PersonEntity: Name: NameFull: Cord, Matthieu – PersonEntity: Name: NameFull: Valle, Eduardo – PersonEntity: Name: NameFull: de A. Araújo, Arnaldo IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 05 Text: May2013 Type: published Y: 2013 Identifiers: – Type: issn-print Value: 10773142 Numbering: – Type: volume Value: 117 – Type: issue Value: 5 Titles: – TitleFull: Computer Vision & Image Understanding Type: main |
| ResultId | 1 |