Learning Where to Attend with Deep Architectures for Image Tracking.

Saved in:
Bibliographic Details
Title: Learning Where to Attend with Deep Architectures for Image Tracking.
Authors: Denil, Misha, Bazzani, Loris, Larochelle, Hugo, de Freitas, Nando
Source: Neural Computation. Aug2012, Vol. 24 Issue 8, p2151-2184. 34p.
Subjects: Machine learning, Computer architecture, Statistical methods in image analysis, Neurosciences, Boltzmann machine, Ranking (Statistics), Performance evaluation
Abstract: We discuss an attentional model for simultaneous object tracking and recognition that is driven by gaze data. Motivated by theories of perception, the model consists of two interacting pathways, identity and control, intended tomirror the what andwhere pathways in neuroscience models. The identity pathway models object appearance and performs classification using deep (factored)-restricted Boltzmann machines. At each point in time, the observations consist of foveated images, with decaying resolution toward the periphery of the gaze. The control pathway models the location, orientation, scale, and speed of the attended object. The posterior distribution of these states is estimated with particle filtering. Deeper in the control pathway, we encounter an attentional mechanism that learns to select gazes so as to minimize tracking uncertainty. Unlike in our previous work, we introduce gaze selection strategies that operate in the presence of partial information and on a continuous action space. We show that a straightforward extension of the existing approach to the partial information setting results in poor performance, and we propose an alternative method based on modeling the reward surface as a gaussian process. This approach gives good performance in the presence of partial information and allows us to expand the action space from a small, discrete set of fixation points to a continuous domain. [ABSTRACT FROM AUTHOR]
Copyright of Neural Computation is the property of MIT Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Psychology and Behavioral Sciences Collection
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: pbh
DbLabel: Psychology and Behavioral Sciences Collection
An: 76625570
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Learning Where to Attend with Deep Architectures for Image Tracking.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Denil%2C+Misha%22">Denil, Misha</searchLink><br /><searchLink fieldCode="AR" term="%22Bazzani%2C+Loris%22">Bazzani, Loris</searchLink><br /><searchLink fieldCode="AR" term="%22Larochelle%2C+Hugo%22">Larochelle, Hugo</searchLink><br /><searchLink fieldCode="AR" term="%22de+Freitas%2C+Nando%22">de Freitas, Nando</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Neural+Computation%22">Neural Computation</searchLink>. Aug2012, Vol. 24 Issue 8, p2151-2184. 34p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+architecture%22">Computer architecture</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+methods+in+image+analysis%22">Statistical methods in image analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Neurosciences%22">Neurosciences</searchLink><br /><searchLink fieldCode="DE" term="%22Boltzmann+machine%22">Boltzmann machine</searchLink><br /><searchLink fieldCode="DE" term="%22Ranking+%28Statistics%29%22">Ranking (Statistics)</searchLink><br /><searchLink fieldCode="DE" term="%22Performance+evaluation%22">Performance evaluation</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: We discuss an attentional model for simultaneous object tracking and recognition that is driven by gaze data. Motivated by theories of perception, the model consists of two interacting pathways, identity and control, intended tomirror the what andwhere pathways in neuroscience models. The identity pathway models object appearance and performs classification using deep (factored)-restricted Boltzmann machines. At each point in time, the observations consist of foveated images, with decaying resolution toward the periphery of the gaze. The control pathway models the location, orientation, scale, and speed of the attended object. The posterior distribution of these states is estimated with particle filtering. Deeper in the control pathway, we encounter an attentional mechanism that learns to select gazes so as to minimize tracking uncertainty. Unlike in our previous work, we introduce gaze selection strategies that operate in the presence of partial information and on a continuous action space. We show that a straightforward extension of the existing approach to the partial information setting results in poor performance, and we propose an alternative method based on modeling the reward surface as a gaussian process. This approach gives good performance in the presence of partial information and allows us to expand the action space from a small, discrete set of fixation points to a continuous domain. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Neural Computation is the property of MIT Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=76625570
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1162/NECO_a_00312
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 34
        StartPage: 2151
    Subjects:
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Computer architecture
        Type: general
      – SubjectFull: Statistical methods in image analysis
        Type: general
      – SubjectFull: Neurosciences
        Type: general
      – SubjectFull: Boltzmann machine
        Type: general
      – SubjectFull: Ranking (Statistics)
        Type: general
      – SubjectFull: Performance evaluation
        Type: general
    Titles:
      – TitleFull: Learning Where to Attend with Deep Architectures for Image Tracking.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Denil, Misha
      – PersonEntity:
          Name:
            NameFull: Bazzani, Loris
      – PersonEntity:
          Name:
            NameFull: Larochelle, Hugo
      – PersonEntity:
          Name:
            NameFull: de Freitas, Nando
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 08
              Text: Aug2012
              Type: published
              Y: 2012
          Identifiers:
            – Type: issn-print
              Value: 08997667
          Numbering:
            – Type: volume
              Value: 24
            – Type: issue
              Value: 8
          Titles:
            – TitleFull: Neural Computation
              Type: main
ResultId 1